VibeVoice-Realtime-0.5B β€” with encoder (voice cloning)

microsoft/VibeVoice-Realtime-0.5B with the missing acoustic encoder added β€” enabling voice cloning from your own audio.

Usage

pip install "transformers==4.51.3" torch soundfile
pip install git+https://github.com/microsoft/VibeVoice

# get the scripts (the model itself downloads automatically on first run)
huggingface-cli download mohammed-bahumaish/vibevoice-realtime-0.5b-with-encoder \
    make_voice_prompt.py run_tts.py --local-dir .

# 1) build a voice prompt from ~15-30s of reference audio
python make_voice_prompt.py \
    --voice_wav my_voice.wav \
    --transcript "exact transcript of the reference audio" \
    --output my_voice.pt

# 2) speak anything in that voice
python run_tts.py \
    --voice_pt my_voice.pt \
    --text "Hello! This works with the stock Microsoft inference code." \
    --output out.wav

The .pt files are drop-in compatible with Microsoft's own demos, like the prebaked demo/voices/streaming_model/*.pt voices.

Tips

  • transformers must be 4.51.x β€” 5.x silently breaks the model.
  • Use a true 24 kHz+ recording, β‰₯ 15 s, clean single speaker.
  • Pass --transcript explicitly for best results (auto-transcription is English-only).

License

MIT. Base model by Microsoft; its responsible-use guidelines apply β€” clone only voices you have the right to use.

Downloads last month
670
Safetensors
Model size
1B params
Tensor type
F32
Β·
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for mohammed-bahumaish/vibevoice-realtime-0.5b-with-encoder

Finetuned
(17)
this model

Space using mohammed-bahumaish/vibevoice-realtime-0.5b-with-encoder 1