Whisper Tiny Dual-Scale Turn Detector

A compact audio-only classifier that predicts whether a speaker has completed their turn or is holding the floor. It is designed to run after a lightweight VAD observes a candidate pause. This model does not require ASR or text at inference time.

This card is generated from the packaged run artifacts. Missing results are shown as Not available instead of being estimated or hand-entered.

Model details

  • Base encoder: openai/whisper-tiny
  • Architecture: dual_scale
  • Input: 16000 Hz mono audio, last 8.0 s
  • Output: calibrated probability of end-of-turn / COMPLETE
  • INT8 ONNX size: 10.16 MiB
  • INT8 method: dynamic_weight_only
  • INT8 maximum probability difference from FP32: 0.017759
  • Export parity target passed: True

The encoder feeds masked global attention pooling and a recent-window attention/mean/max branch. The fused representation predicts turn completion. Separate mid-filler and end-filler auxiliary outputs supervise training but are not required by the exported inference graph.

Test results

The candidate temperature and endpoint policy are selected on validation data and frozen before test evaluation.

Metric Test value
Examples 21,995
F1 0.7399
Balanced accuracy 0.8248
AUROC 0.9361
Average precision 0.7983
False-cutoff rate 0.0472
False-hold rate 0.3033
Expected calibration error 0.1043

Causal turn policy

  • Threshold: 0.3800
  • Temperature: 2.5522
  • Minimum silence: 300 ms
  • Fallback timeout: 1,000 ms
  • Turn-level false-cutoff rate: 0.0497
  • Mean endpoint latency: 512.3 ms
  • P95 endpoint latency: 1000.0 ms
  • Calibration split: validation
  • Test policy tuning performed: False

Important slices

Slice Count F1 False-cutoff rate
language:hin 3,124 0.7656 0.1144
language:eng 18,871 0.7340 0.0362
filler:mid 5,099 0.7239 0.0840
filler:end 2,305 Not available Not available
filler:any 6,235 0.7034 0.0773
hard:incomplete_filler 5,267 Not available Not available
kind:causal_internal_pause 13,040 Not available Not available
kind:original 8,955 0.7905 0.0648

Robustness

The robustness subset is deterministically stratified across language, label, original/causal examples, filler type, and synthetic status. Each corruption uses the same selected records.

Condition F1 False-cutoff rate
clean 0.7044 0.0799
clipping 0.6835 0.0839
gain_low 0.6904 0.0852
mulaw 0.7042 0.0826
noise_10db 0.5283 0.1358
noise_20db 0.6905 0.0826
noise_5db 0.3679 0.1292
reverb 0.5861 0.1438
speed_0.9 0.6653 0.1012
speed_1.1 0.6608 0.0759
telephone 0.6477 0.1105

CPU latency

  • Model-only P50/P95/P99: 39.15 / 75.89 / 79.14 ms
  • End-to-end P50/P95/P99: 47.36 / 82.36 / 87.90 ms
  • Runtime: Linux-6.8.0-124-generic-x86_64-with-glibc2.35, 128 logical CPUs

End-to-end latency includes waveform standardization, log-Mel extraction, probability calibration, and ONNX inference. It excludes audio decoding/disk I/O and the configured silence wait.

Public baseline

Smart Turn v3.2 is pinned by revision. Its temperature, threshold, action delay, and timeout are selected on validation data under the same configured false-cutoff budget as the candidate, then both policies are frozen for the paired test comparison.

Model Test F1
Candidate 0.7399
Smart Turn v3.2 0.4364
  • Candidate minus Smart Turn F1: 0.3035
  • 95% group-bootstrap interval: [0.2884, 0.3182]
  • Matched validation false-cutoff budget: 0.0500

Training and evaluation data

Training and in-domain evaluation use only the Hindi (hin) and English (eng) subsets of the provided Smart Turn v3.2 train/test dataset family. endpoint_bool supplies the main completion label; midfiller and endfiller supply auxiliary supervision. Causal internal-pause examples and mined hard negatives teach the model not to interrupt a speaker who is likely to continue.

The test split remains separate from threshold, temperature, policy, and baseline selection. Reported confidence intervals use parent-turn group bootstrapping. Robustness results cover telephone filtering, mu-law, additive noise, speed changes, low gain, clipping, and reverb.

Limitations and honest scope

  • The source has Hindi/English metadata but no human-verified Hinglish or code-switch label. Consequently, this release does not claim a measured Hinglish-specific test score. Hindi and filler-focused results are relevant proxies, not a Hinglish ground-truth benchmark.
  • English rows are not guaranteed to be exclusively Indian English.
  • Some source audio is synthetic, and the source does not provide reliable speaker identities.
  • The detector uses audio only; it cannot use conversation history, transcript semantics, gaze, or dialog state.
  • It is not a backchannel, barge-in, speaker-diarization, or safety classifier.
  • Review the upstream dataset terms before distributing derived weights. Source audio is not included in this package.

Minimal inference

from turn_detector.inference import TurnDetector
from turn_detector.audio import load_audio

detector = TurnDetector("hinglish-turn.int8.onnx")
audio, sample_rate = load_audio("candidate_pause.wav")
prediction = detector.score(audio, sample_rate)
print(prediction.probability, prediction.decision)

See evaluation/evaluation_report.json, evaluation/baselines/baseline_report.json, calibration_report.json, and export_report.json for the complete machine-readable results.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
8.3M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Space using Mayank022/hinglish-turn-detector-whisper-tiny-dual-scale 1