Audio8 ASR Infinite: rolling KV cache enables 24/7 drift-free streaming ASR (Apache 2.0)

Edge0 open-sources Audio8 ASR Infinite, a 4B-parameter streaming speech recognition model (Apache 2.0, Chinese and English). Core: a rolling KV cache with a fixed 30-second context window plus exact RoPE re-basing keeps positions mathematically consistent as old frames are evicted - memory and latency stay constant for unlimited-length 24/7 transcription with no drift. The native streaming architecture decodes up to 12.5 times per second, with selectable audio clocks of 80/120/160 ms (12.5/8.3/6.25 decisions per second) and configurable transcription delay of 240-560 ms from a single checkpoint. Semantic VAD distinguishes thinking pauses, stuttering and real end of turn where traditional acoustic VAD fails. Benchmarks: aishell1/test CER 1.750% and aishell4 2.893%, far below Voxtral-Mini-4B-Realtime (16.8/16.5) and nemotron-3.5-asr-streaming (12.9/14.7); slightly behind on English LibriSpeech (test.clean 3.04 vs 2.21). Ships via transformers and an adapted vLLM build (Docker), targeting always-on microphone scenarios like meeting transcription, call analytics and full-duplex voice agents; AK also built a Hugging Face Space demo. Directly usable as the real-time transcription layer for robot voice interaction.

![Fish Audio upgrades its ASR model: speaker identification and inline emotion cues such as [laughter]](/static/img/twitter/FishAudio_2104617131920748857.card.webp)



