Audio-Text-to-Text
Transformers
Safetensors
English
Chinese
qwen2
text-generation
speech-language-model
streaming
audio
multimodal
qwen2.5-omni
text-generation-inference
Instructions to use zhifeixie/AudioInteraction with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use zhifeixie/AudioInteraction with Transformers:
# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("zhifeixie/AudioInteraction") model = AutoModelForCausalLM.from_pretrained("zhifeixie/AudioInteraction", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 5,937 Bytes
49707c0 7fda347 c1144d4 c0caf87 7fda347 c1144d4 49707c0 7fda347 e0c9e81 c1144d4 7fda347 c1144d4 7fda347 c1144d4 7fda347 c1144d4 e0c9e81 c1144d4 e0c9e81 c0caf87 c1144d4 7fda347 c1144d4 7fda347 e0c9e81 c1144d4 e0c9e81 c0caf87 c1144d4 e0c9e81 c1144d4 7fda347 c0caf87 c1144d4 c0caf87 c1144d4 7fda347 c1144d4 7fda347 c1144d4 e0c9e81 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 | ---
datasets:
- zhifeixie/StreamAudio-2M
language:
- en
- zh
library_name: transformers
license: apache-2.0
pipeline_tag: audio-text-to-text
tags:
- speech-language-model
- streaming
- audio
- multimodal
- qwen2.5-omni
---
# Audio-Interaction: Streaming Audio-In, Text-Out Conversational Model
[**Project Page**](https://xzf-thu.github.io/Audio-Interaction/) | [**Code**](https://github.com/xzf-thu/Audio-Interaction) | [**Model**](https://huggingface.co/zhifeixie/Audio-Interaction) | [**Dataset**](https://huggingface.co/datasets/zhifeixie/StreamAudio-2M) | [**Paper**](https://huggingface.co/papers/2606.05121)
Audio-Interaction is a unified streaming model that listens to audio in real time and decides, at each audio chunk, whether to keep listening or to start replying with text. It formalizes the "perceive-decide-respond" loop, allowing the model to handle conventional offline tasks (ASR, S2TT) while adding online capabilities like proactive intervention and real-time voice chatting.
The model alternates between a **LISTENING** state, where it consumes one encoder-output chunk per step and emits either `KEEP_SILENCE` or `TEXT_BEGIN`, and a **SPEAKING** state, where it autoregressively generates a text turn until `TEXT_END` and then returns to listening for the next chunk.
## Model Details
- **Model name:** Audio-Interaction
- **Task:** Streaming audio-conditioned text generation (audio in, text out)
- **Audio encoder:** Qwen2.5-Omni audio tower (chunk-wise)
- **Audio framing:** 16 kHz, padded to 0.4-second (6400-sample) boundaries; 10 encoder-output frames per chunk
- **Decoding states:** LISTENING (emits `KEEP_SILENCE` / `TEXT_BEGIN`) and SPEAKING (emits text until `TEXT_END`)
- **Default sampling:** temperature 0.3, top-k 3
- **Default max new tokens:** 4096 per session
- **License:** Apache-2.0
## Repository Contents
```text
Audio-Interaction/
βββ model-00001-of-00004.safetensors # LM weights, sharded (β4 GB each)
βββ model-00002-of-00004.safetensors
βββ model-00003-of-00004.safetensors
βββ model-00004-of-00004.safetensors
βββ model.safetensors.index.json # Shard index consumed by safetensors loader
βββ config.json # Top-level model config
βββ generation_config.json # Generation defaults
βββ model_config.yaml # GPT config consumed by Config.from_file
βββ hyperparameters.yaml # Training-time hyperparameters (reference)
βββ tokenizer.json # Tokenizer
βββ tokenizer_config.json
βββ MiniOmni3_ChunkwisedEncoder.pth # Audio encoder weights (Qwen2.5-Omni audio tower)
βββ qwen25OmniConfig/ # Audio-encoder config (nested: thinker_config.audio_config)
```
## Intended Use
Audio-Interaction is intended for streaming conversational agents that need to react to audio as it arrives β for example, voice assistants that may interject mid-utterance, alarms that respond to ambient sound, or low-latency dialogue systems where waiting for a full utterance before replying is too slow.
## Quick Start
### Installation
```bash
git clone https://github.com/xzf-thu/Audio-Interaction.git
cd Audio-Interaction
conda create -n Audio-Interaction python=3.10 -y
conda activate Audio-Interaction
pip install -r requirements.txt
```
### Download the checkpoint
From the `Audio-Interaction` project root, pull the weights into `checkpoints/`:
```python
from huggingface_hub import snapshot_download
snapshot_download(repo_id="zhifeixie/Audio-Interaction", local_dir="checkpoints")
```
`snapshot_download` is the recommended path β it pulls every file and resumes on interruption.
### Python Usage
```python
from src.miniomni3.generate.run import run_inference
run_inference(
checkpoint_dir="checkpoints",
audio_paths=["/path/to/audio.wav"], # offline mode: one round per path
device="cuda:0", # or "mps" / "cpu"
)
```
## Streaming Protocol
A single session looks like:
```text
[system prompt tokens]
ββββ LISTENING ββββ
β AUDIO_BEGIN PAD*10 ASSISTANT β KEEP_SILENCE (keep listening)
β AUDIO_BEGIN PAD*10 ASSISTANT β TEXT_BEGIN EMOTION (start replying)
βββββββββββββββββββ
ββββ SPEAKING βββββ
β β¦ text tokens β¦ TEXT_END (reply finished)
βββββββββββββββββββ
ββββ LISTENING ββββ (next audio chunk)
β¦
```
The model is trained to emit at most one `TEXT_BEGIN` per audio chunk. Each assistant turn begins with `TEXT_BEGIN`, followed by an emotion token, the reply tokens, and `TEXT_END`. Turns starting with `KEEP_SILENCE` indicate the model chose not to respond to that chunk.
## Limitations
- The model produces text, not speech. Pair it with a TTS system for end-to-end voice interaction.
- Audio must be 16 kHz mono; non-conforming inputs are resampled and padded to 0.4-second boundaries.
- Decisions are made at 0.4-second granularity (one encoder chunk), which sets a floor on response-onset latency.
- Trailing partial audio chunks shorter than 10 encoder frames are dropped before generation.
## Citation
```bibtex
@misc{xie2026audiointeractionmodel,
title={Audio Interaction Model},
author={Zhifei Xie and Zihang Liu and Ze An and Xiaobin Hu and Yue Liao and Ziyang Ma and Dongchao Yang and Mingbao Lin and Deheng Ye and Shuicheng Yan and Chunyan Miao},
year={2026},
eprint={2606.05121},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2606.05121},
}
```
## Acknowledgements
Audio-Interaction builds on the Qwen2.5-Omni audio encoder. We thank the Qwen team and the maintainers of OpenAI Whisper for the audio-loading utilities used in this project. |