Attuned Resonance β€” Voice Tone Classifier

The 4th model in the Attuned Resonance call-center cascade. Detects caller emotion across 10 classes from short audio clips (2–8 seconds).

Status: v1 (frozen-backbone baseline). Full fine-tune planned for v2.2. Hopeful and urgent classes are undersampled β€” backfill pending. Part of: Attuned Resonance β€” a 4-model cascade for inbound call routing.

Model Description

Architecture facebook/wav2vec2-base (frozen, 95M params) β†’ mean-pool over hidden frames (masked) β†’ 768β†’256 GELU+dropout(0.3)β†’10 MLP head
Training data 6,021 synthetic clips generated by ElevenLabs eleven_v3 with audio-tag labels ([angry], [hopeful], etc.)
Sample rate 16 kHz mono (resampled from 44.1 kHz mp3)
Output softmax over EMOTION_LABELS (10 classes), with predict() returning the top label, confidence, and full probability distribution

Emotion taxonomy

EMOTION_LABELS = [
    "angry", "frustrated", "sad", "calm", "anxious",
    "satisfied", "confused", "neutral", "urgent", "hopeful",
]

Each label is mapped to one or more ElevenLabs v3 audio tags during synthesis. The synthesis instruction is the ground-truth label β€” no manual annotation. See models/voice_tone/labels.py for the tag map.

Intended use

  • Cascade integration: feed the 10-class probability vector to the Outcome Predictor (planned v2.2 fusion).
  • Standalone emotion detection on short call-center audio.
  • Research and benchmarking on synthesis-as-label vs human-labeled emotion datasets.

Out-of-scope

  • Production routing decisions without human-in-the-loop review.
  • Languages other than English.
  • Long-form audio (clips > ~30 seconds β€” the mean-pool collapses too much).
  • Real spontaneous speech until v2.2 fine-tune validates real-world transfer.

Training details

Hyperparameter Value
Backbone facebook/wav2vec2-base (frozen β€” requires_grad=False for all backbone params)
Head 768 β†’ 256 (GELU + dropout 0.3) β†’ 10
Optimizer AdamW, lr=1e-3, weight_decay=1e-4
Loss Cross-entropy
Sampler WeightedRandomSampler (compensates for hopeful 4.3Γ— imbalance)
Epochs 20
Batch size 64
Split Voice-stratified: 26 train / 4 val / 4 test voices (deterministic seed=42)
Hardware CPU (Intel, 12 cores). Frozen backbone caches features once β†’ linear-head training is seconds.

Evaluation

Held-out test split: 4 voices the model never saw during training. The voice-stratified split is the right benchmark for an emotion classifier β€” it forces the model to generalize across speakers rather than lean on per-voice timbre cues.

Metric Value
Best val accuracy 64.7%
Test accuracy 59.9%
Random baseline (10-class) 10.0%

Per-class test accuracy

Class Acc n Notes
frustrated 91.2% 80 Distinctive prosody β€” easy
angry 85.0% 80 High arousal β€” easy
urgent 77.8% 72 Distinctive β€” easy
hopeful 73.7% 19 Small test n β€” high variance
neutral 67.5% 80 Tag-free baseline
satisfied 56.2% 80
sad 50.0% 80
calm 38.8% 80 Confused with anxious / neutral
anxious 36.2% 80 Confused with confused / sad
confused 35.0% 80 Confused with anxious / calm

The high/low split is interpretable: extreme arousal states (angry, frustrated, urgent) have distinctive prosody and are easy. Mid-arousal states (calm, anxious, confused) share acoustic features in the frozen wav2vec2 representation. A full fine-tune (v2.2) is the obvious lever.

Limitations

  1. Synthesis-as-label prior. The classifier learns "what eleven_v3 thinks [angry] sounds like." Real callers do not speak with explicit synthesis tags; their prosody distribution may differ. Real-audio held-out evaluation is the v2.2 milestone.
  2. Class imbalance. Hopeful: 154 clips (vs 660 for most classes). Test set has only 19 hopeful examples β€” metrics for that class are noisy.
  3. 34 voices, English only. No accent / age / gender audit. ElevenLabs voices skew accent-neutral.
  4. Short clips only. Mean-pooled embedding collapses long temporal structure. For long-form calls, a sliding-window approach is needed.
  5. Frozen backbone ceiling. The backbone wasn't trained on emotion-discriminative objectives; mid-arousal class confusion reflects that.

Usage

from models.voice_tone.inference import VoiceToneClassifier

clf = VoiceToneClassifier.load('trained_models/voice_tone')
result = clf.predict('path/to/clip.mp3')

# result.emotion       -> 'angry'
# result.confidence    -> 0.953
# result.probs         -> {'angry': 0.953, 'urgent': 0.022, ...}

Files

  • head.safetensors β€” trained head weights (256-hidden MLP).
  • config.json β€” architecture config + emotion labels.
  • metrics.json β€” full training history + per-class test metrics.

The wav2vec2 backbone is loaded from facebook/wav2vec2-base at inference time; we do not redistribute it.

Citation

@misc{rubin2026attunedresonance,
  title  = {Attuned Resonance: A Multi-Model Cascade for Inbound Call-Center Routing},
  author = {Theodore Rubin},
  year   = {2026},
  url    = {https://github.com/tedrubin80/CEPM}
}

License

CC-BY-NC-4.0 (research/educational use). Not a production system.

Downloads last month
11
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for datamatters24/attuned-resonance-voice-tone

Finetuned
(1009)
this model