Whisper Tiny Dual-Scale Turn Detector
A compact audio-only classifier that predicts whether a speaker has completed their turn or is holding the floor. It is designed to run after a lightweight VAD observes a candidate pause. This model does not require ASR or text at inference time.
This card is generated from the packaged run artifacts. Missing results are shown as Not available instead of being estimated or hand-entered.
Model details
- Base encoder:
openai/whisper-tiny - Architecture:
dual_scale - Input: 16000 Hz mono audio, last 8.0 s
- Output: calibrated probability of end-of-turn /
COMPLETE - INT8 ONNX size: 10.16 MiB
- INT8 method:
dynamic_weight_only - INT8 maximum probability difference from FP32: 0.017759
- Export parity target passed:
True
The encoder feeds masked global attention pooling and a recent-window attention/mean/max branch. The fused representation predicts turn completion. Separate mid-filler and end-filler auxiliary outputs supervise training but are not required by the exported inference graph.
Test results
The candidate temperature and endpoint policy are selected on validation data and frozen before test evaluation.
| Metric | Test value |
|---|---|
| Examples | 21,995 |
| F1 | 0.7399 |
| Balanced accuracy | 0.8248 |
| AUROC | 0.9361 |
| Average precision | 0.7983 |
| False-cutoff rate | 0.0472 |
| False-hold rate | 0.3033 |
| Expected calibration error | 0.1043 |
Causal turn policy
- Threshold: 0.3800
- Temperature: 2.5522
- Minimum silence: 300 ms
- Fallback timeout: 1,000 ms
- Turn-level false-cutoff rate: 0.0497
- Mean endpoint latency: 512.3 ms
- P95 endpoint latency: 1000.0 ms
- Calibration split:
validation - Test policy tuning performed:
False
Important slices
| Slice | Count | F1 | False-cutoff rate |
|---|---|---|---|
language:hin |
3,124 | 0.7656 | 0.1144 |
language:eng |
18,871 | 0.7340 | 0.0362 |
filler:mid |
5,099 | 0.7239 | 0.0840 |
filler:end |
2,305 | Not available | Not available |
filler:any |
6,235 | 0.7034 | 0.0773 |
hard:incomplete_filler |
5,267 | Not available | Not available |
kind:causal_internal_pause |
13,040 | Not available | Not available |
kind:original |
8,955 | 0.7905 | 0.0648 |
Robustness
The robustness subset is deterministically stratified across language, label, original/causal examples, filler type, and synthetic status. Each corruption uses the same selected records.
| Condition | F1 | False-cutoff rate |
|---|---|---|
clean |
0.7044 | 0.0799 |
clipping |
0.6835 | 0.0839 |
gain_low |
0.6904 | 0.0852 |
mulaw |
0.7042 | 0.0826 |
noise_10db |
0.5283 | 0.1358 |
noise_20db |
0.6905 | 0.0826 |
noise_5db |
0.3679 | 0.1292 |
reverb |
0.5861 | 0.1438 |
speed_0.9 |
0.6653 | 0.1012 |
speed_1.1 |
0.6608 | 0.0759 |
telephone |
0.6477 | 0.1105 |
CPU latency
- Model-only P50/P95/P99: 39.15 / 75.89 / 79.14 ms
- End-to-end P50/P95/P99: 47.36 / 82.36 / 87.90 ms
- Runtime:
Linux-6.8.0-124-generic-x86_64-with-glibc2.35, 128 logical CPUs
End-to-end latency includes waveform standardization, log-Mel extraction, probability calibration, and ONNX inference. It excludes audio decoding/disk I/O and the configured silence wait.
Public baseline
Smart Turn v3.2 is pinned by revision. Its temperature, threshold, action delay, and timeout are selected on validation data under the same configured false-cutoff budget as the candidate, then both policies are frozen for the paired test comparison.
| Model | Test F1 |
|---|---|
| Candidate | 0.7399 |
| Smart Turn v3.2 | 0.4364 |
- Candidate minus Smart Turn F1: 0.3035
- 95% group-bootstrap interval: [0.2884, 0.3182]
- Matched validation false-cutoff budget: 0.0500
Training and evaluation data
Training and in-domain evaluation use only the Hindi (hin) and English (eng) subsets of the
provided Smart Turn v3.2 train/test dataset family. endpoint_bool supplies the main completion
label; midfiller and endfiller supply auxiliary supervision. Causal internal-pause examples and
mined hard negatives teach the model not to interrupt a speaker who is likely to continue.
The test split remains separate from threshold, temperature, policy, and baseline selection. Reported confidence intervals use parent-turn group bootstrapping. Robustness results cover telephone filtering, mu-law, additive noise, speed changes, low gain, clipping, and reverb.
Limitations and honest scope
- The source has Hindi/English metadata but no human-verified Hinglish or code-switch label. Consequently, this release does not claim a measured Hinglish-specific test score. Hindi and filler-focused results are relevant proxies, not a Hinglish ground-truth benchmark.
- English rows are not guaranteed to be exclusively Indian English.
- Some source audio is synthetic, and the source does not provide reliable speaker identities.
- The detector uses audio only; it cannot use conversation history, transcript semantics, gaze, or dialog state.
- It is not a backchannel, barge-in, speaker-diarization, or safety classifier.
- Review the upstream dataset terms before distributing derived weights. Source audio is not included in this package.
Minimal inference
from turn_detector.inference import TurnDetector
from turn_detector.audio import load_audio
detector = TurnDetector("hinglish-turn.int8.onnx")
audio, sample_rate = load_audio("candidate_pause.wav")
prediction = detector.score(audio, sample_rate)
print(prediction.probability, prediction.decision)
See evaluation/evaluation_report.json, evaluation/baselines/baseline_report.json,
calibration_report.json, and export_report.json for the complete machine-readable results.