google/fleurs
Viewer • Updated • 768k • 88.6k • 431
How to use deepdml/whisper-tiny-es-mix-norm with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("automatic-speech-recognition", model="deepdml/whisper-tiny-es-mix-norm") # Load model directly
from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq
processor = AutoProcessor.from_pretrained("deepdml/whisper-tiny-es-mix-norm")
model = AutoModelForSpeechSeq2Seq.from_pretrained("deepdml/whisper-tiny-es-mix-norm", device_map="auto")This model is a fine-tuned version of openai/whisper-tiny on the Common Voice 17.0 dataset. It achieves the following results on the evaluation set:
More information needed
More information needed
More information needed
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | Wer Raw | Cer Raw | Wer | Cer |
|---|---|---|---|---|---|---|---|
| 0.5259 | 0.05 | 1000 | 0.5153 | 27.2204 | 9.7771 | 27.1670 | 9.7663 |
| 0.4068 | 0.1 | 2000 | 0.4695 | 24.9782 | 9.1500 | 24.9598 | 9.1466 |
| 0.2546 | 0.15 | 3000 | 0.4376 | 23.3862 | 8.6061 | 23.3767 | 8.6047 |
| 0.2461 | 0.2 | 4000 | 0.4191 | 22.2918 | 8.0878 | 22.2880 | 8.0871 |
| 0.2203 | 0.25 | 5000 | 0.4099 | 22.2696 | 8.1756 | 22.2658 | 8.1750 |
| 0.2441 | 0.3 | 6000 | 0.4001 | 21.4796 | 7.8921 | 21.4790 | 7.8920 |
| 0.2318 | 0.35 | 7000 | 0.3908 | 21.4504 | 7.9370 | 21.4504 | 7.9370 |
| 0.4077 | 0.4 | 8000 | 0.3833 | 20.7589 | 7.6047 | 20.7589 | 7.6047 |
| 0.1844 | 0.45 | 9000 | 0.3808 | 20.4431 | 7.5154 | 20.4431 | 7.5154 |
| 0.2673 | 0.5 | 10000 | 0.3750 | 20.3490 | 7.3850 | 20.3484 | 7.3849 |
| 0.1677 | 0.55 | 11000 | 0.3726 | 20.3262 | 7.5745 | 20.3262 | 7.5745 |
| 0.1542 | 0.6 | 12000 | 0.3705 | 19.9080 | 7.2882 | 19.9080 | 7.2882 |
| 0.1609 | 0.65 | 13000 | 0.3647 | 19.8158 | 7.3612 | 19.8158 | 7.3612 |
| 0.1483 | 0.7 | 14000 | 0.3607 | 19.8438 | 7.4117 | 19.8438 | 7.4117 |
| 0.1343 | 0.75 | 15000 | 0.3607 | 19.5616 | 7.2026 | 19.5616 | 7.2026 |
| 0.1379 | 1.0116 | 16000 | 0.3598 | 19.5146 | 7.1944 | 19.5146 | 7.1944 |
| 0.1271 | 1.0616 | 17000 | 0.3589 | 19.7453 | 7.2933 | 19.7453 | 7.2933 |
| 0.1309 | 1.1117 | 18000 | 0.3558 | 19.7154 | 7.3180 | 19.7154 | 7.3180 |
| 0.1561 | 1.1617 | 19000 | 0.3573 | 19.5197 | 7.0518 | 19.5197 | 7.0518 |
| 0.206 | 1.2117 | 20000 | 0.3560 | 19.6036 | 7.2565 | 19.6036 | 7.2565 |
Please cite the model using the following BibTeX entry:
@misc{deepdml/whisper-tiny-es-mix-norm,
title={Fine-tuned Whisper tiny ASR model for speech recognition in Spanish},
author={Jimenez, David},
howpublished={\url{https://huggingface.co/deepdml/whisper-tiny-es-mix-norm}},
year={2026}
}