Instructions to use datamatters24/attuned-resonance-voice-tone with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use datamatters24/attuned-resonance-voice-tone with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("audio-classification", model="datamatters24/attuned-resonance-voice-tone")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("datamatters24/attuned-resonance-voice-tone", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Attuned Resonance β Voice Tone Classifier
The 4th model in the Attuned Resonance call-center cascade. Detects caller emotion across 10 classes from short audio clips (2β8 seconds).
Status: v1 (frozen-backbone baseline). Full fine-tune planned for v2.2. Hopeful and urgent classes are undersampled β backfill pending. Part of: Attuned Resonance β a 4-model cascade for inbound call routing.
Model Description
| Architecture | facebook/wav2vec2-base (frozen, 95M params) β mean-pool over hidden frames (masked) β 768β256 GELU+dropout(0.3)β10 MLP head |
| Training data | 6,021 synthetic clips generated by ElevenLabs eleven_v3 with audio-tag labels ([angry], [hopeful], etc.) |
| Sample rate | 16 kHz mono (resampled from 44.1 kHz mp3) |
| Output | softmax over EMOTION_LABELS (10 classes), with predict() returning the top label, confidence, and full probability distribution |
Emotion taxonomy
EMOTION_LABELS = [
"angry", "frustrated", "sad", "calm", "anxious",
"satisfied", "confused", "neutral", "urgent", "hopeful",
]
Each label is mapped to one or more ElevenLabs v3 audio tags during synthesis. The synthesis instruction is the ground-truth label β no manual annotation. See models/voice_tone/labels.py for the tag map.
Intended use
- Cascade integration: feed the 10-class probability vector to the Outcome Predictor (planned v2.2 fusion).
- Standalone emotion detection on short call-center audio.
- Research and benchmarking on synthesis-as-label vs human-labeled emotion datasets.
Out-of-scope
- Production routing decisions without human-in-the-loop review.
- Languages other than English.
- Long-form audio (clips > ~30 seconds β the mean-pool collapses too much).
- Real spontaneous speech until v2.2 fine-tune validates real-world transfer.
Training details
| Hyperparameter | Value |
|---|---|
| Backbone | facebook/wav2vec2-base (frozen β requires_grad=False for all backbone params) |
| Head | 768 β 256 (GELU + dropout 0.3) β 10 |
| Optimizer | AdamW, lr=1e-3, weight_decay=1e-4 |
| Loss | Cross-entropy |
| Sampler | WeightedRandomSampler (compensates for hopeful 4.3Γ imbalance) |
| Epochs | 20 |
| Batch size | 64 |
| Split | Voice-stratified: 26 train / 4 val / 4 test voices (deterministic seed=42) |
| Hardware | CPU (Intel, 12 cores). Frozen backbone caches features once β linear-head training is seconds. |
Evaluation
Held-out test split: 4 voices the model never saw during training. The voice-stratified split is the right benchmark for an emotion classifier β it forces the model to generalize across speakers rather than lean on per-voice timbre cues.
| Metric | Value |
|---|---|
| Best val accuracy | 64.7% |
| Test accuracy | 59.9% |
| Random baseline (10-class) | 10.0% |
Per-class test accuracy
| Class | Acc | n | Notes |
|---|---|---|---|
| frustrated | 91.2% | 80 | Distinctive prosody β easy |
| angry | 85.0% | 80 | High arousal β easy |
| urgent | 77.8% | 72 | Distinctive β easy |
| hopeful | 73.7% | 19 | Small test n β high variance |
| neutral | 67.5% | 80 | Tag-free baseline |
| satisfied | 56.2% | 80 | |
| sad | 50.0% | 80 | |
| calm | 38.8% | 80 | Confused with anxious / neutral |
| anxious | 36.2% | 80 | Confused with confused / sad |
| confused | 35.0% | 80 | Confused with anxious / calm |
The high/low split is interpretable: extreme arousal states (angry, frustrated, urgent) have distinctive prosody and are easy. Mid-arousal states (calm, anxious, confused) share acoustic features in the frozen wav2vec2 representation. A full fine-tune (v2.2) is the obvious lever.
Limitations
- Synthesis-as-label prior. The classifier learns "what
eleven_v3thinks[angry]sounds like." Real callers do not speak with explicit synthesis tags; their prosody distribution may differ. Real-audio held-out evaluation is the v2.2 milestone. - Class imbalance. Hopeful: 154 clips (vs 660 for most classes). Test set has only 19 hopeful examples β metrics for that class are noisy.
- 34 voices, English only. No accent / age / gender audit. ElevenLabs voices skew accent-neutral.
- Short clips only. Mean-pooled embedding collapses long temporal structure. For long-form calls, a sliding-window approach is needed.
- Frozen backbone ceiling. The backbone wasn't trained on emotion-discriminative objectives; mid-arousal class confusion reflects that.
Usage
from models.voice_tone.inference import VoiceToneClassifier
clf = VoiceToneClassifier.load('trained_models/voice_tone')
result = clf.predict('path/to/clip.mp3')
# result.emotion -> 'angry'
# result.confidence -> 0.953
# result.probs -> {'angry': 0.953, 'urgent': 0.022, ...}
Files
head.safetensorsβ trained head weights (256-hidden MLP).config.jsonβ architecture config + emotion labels.metrics.jsonβ full training history + per-class test metrics.
The wav2vec2 backbone is loaded from facebook/wav2vec2-base at inference time; we do not redistribute it.
Citation
@misc{rubin2026attunedresonance,
title = {Attuned Resonance: A Multi-Model Cascade for Inbound Call-Center Routing},
author = {Theodore Rubin},
year = {2026},
url = {https://github.com/tedrubin80/CEPM}
}
License
CC-BY-NC-4.0 (research/educational use). Not a production system.
- Downloads last month
- 11
Model tree for datamatters24/attuned-resonance-voice-tone
Base model
facebook/wav2vec2-base