VibeVoice: a 7.5Hz token rate buys 90 minutes of speech in one pass
vibevoice
The open frontier speech model family from Microsoft, repository microsoft/VibeVoice (MIT, 54,531 stars, verified through the GitHub API). We catalogue it as one asset rather than three because TTS, Realtime and ASR share a single technical core: both the acoustic and the semantic tokenizer are continuous rather than quantised into discrete codebooks, and the frame rate is pushed down to 7.5Hz. That 7.5Hz is where every capability comes from. Ninety minutes of audio at 25Hz is 135,000 tokens and does not fit a 64K context, while at 7.5Hz it is 40,500, so single pass synthesis of 90 minutes and single pass transcription of 60 minutes are two directions of the same fact rather than two separate engineering feats. Five product lines each carry their own ceiling: TTS-1.5B at 90 minutes per run with up to four speakers, Realtime-0.5B at roughly 300ms to first packet, ASR-7B transcribing 60 minutes in one pass while emitting who said what and when with custom hotwords across more than 50 languages, ASR-Streaming emitting text as speech arrives, and ASR-BitNet using heterogeneous quantisation to compress 4.62GB into 1.58GB with RTF under 1 on three or more CPU threads and no GPU. The ASR line is the one worth remembering: conventional long form transcription chains slicing, recognition, diarisation and timestamp alignment, and the global context is lost at the slice boundary, while VibeVoice-ASR fuses recognition, diarisation and timestamps into a single generation, which is the difference between usable and unusable for meeting minutes and support QA. The architecture is next token diffusion, with the LLM owning conversational direction and the diffusion head owning acoustic detail, because turn consistency over a long dialogue is a language problem and only the LLM can hold four speakers distinct across 90 minutes. One fact must be recorded plainly: on 2025-09-05 the team removed the VibeVoice-TTS code from the repository, stating that usage inconsistent with its declared intent had been observed; the weights remain on Hugging Face but the Quick Try section reads Disabled, the official inference scripts are gone and today the TTS path runs through community reproductions, with the model page advising against commercial or real world use without further testing. Every public move since has been on the ASR side, so the commercially usable half is ASR rather than TTS. Its position in tables published by others must also be reported as written: the CosyVoice 3 README lists it at test-zh CER 1.16 with similarity 74.4 and test-en WER 3.04 with similarity 68.9, against CosyVoice3 base at 1.21 and 78.0. Short clip zero shot cloning is not its lane; the three comparisons that matter are how long a single run can be, whether four speakers stay distinct, and whether one hour of meeting audio can be processed in one pass, and on those three the open source field has almost no rival. Boundaries: no official TTS inference path; MIT covers code but not the usage limits on the model page; long form failure is global, since a 90 minute single pass has no retry one segment at a time fallback; the official Risks and Limitations section names deepfakes and requires disclosure of AI involvement. Graded B (confirmed): what is confirmed is the frame rate arithmetic, the published ceilings of each line, the code removal and the licence boundary, not audio quality, which we only relay without recomputing or blind listening.
- CONFIDENCE
- Confirmed
- Two or more independent sources, or reproduced by our harness
- KEY METRIC
- 单次长音频上限(TTS 4 说话人 / ASR 单遍)
- Confirmed · 2026-09
- MATURITY
- Research
- research → demo → product → production
Our takeWe grade it B (confirmed). The basis is not that we reproduced a 90 minute run but that the claim carries three independent sources of evidence: the VibeVoice-TTS technical report was accepted as an Oral at ICLR 2026 (third party review), VibeVoice-ASR entered Azure AI Foundry Labs on 2026-03-12 (a first party managed product surface rather than a demo page), and it shipped as part of an official Hugging Face Transformers release on 2026-03-06 (third party library integration). The repository itself stands at 54,531 stars under MIT with the last push on 2026-09-03, verified by us through the GitHub API. We have never run a 90 minute synthesis and have not recomputed the ASR DER or cpWER numbers, so grade A is out of reach.
It earns a top row in the audio domain for a reason that has nothing to do with leaderboards: it took the long-form axis that everyone else routes around and made it product grade. Most TTS systems cap a single generation at tens of seconds to a few minutes, and long content is produced by slicing and stitching, which is where timbre drift, prosody breaks and speaker confusion appear. VibeVoice attacks the frame rate instead: both the acoustic and the semantic tokenizers are continuous and run at 7.5 Hz, so 90 minutes of audio is roughly forty thousand tokens (the same content at 25 Hz is 135,000 and at 50 Hz is 270,000), and on the ASR side 60 minutes of audio fits inside a 64K token context in a single pass. That is not an engineering trick but a first principles conversion of context budget into duration, and it is the physical precondition for synthesising four speakers over 90 minutes in one pass or transcribing an hour long meeting into who said what and when.
The most important thing to know when choosing it is that its short-form numbers are not leading. FunAudioLLM lists it in the CosyVoice 3 evaluation table: test-zh CER 1.16 with SS 74.4, test-en WER 3.04 with SS 68.9, and VibeVoice-Realtime at test-en WER 2.05 with SS 63.3. Similarity is below CosyVoice3 at 78.0 and 71.8 and the English error rate is clearly higher. That table comes from a competitor and its direction is unfavourable to VibeVoice, which makes it more credible rather than less. The conclusion is clean: for tens of seconds of high fidelity voice cloning, do not pick it; for multi-speaker podcasts, audiobooks and hour long meeting transcripts, it is close to alone in the open tier.
The second boundary has to be written plainly: on 2025-09-05 Microsoft removed the VibeVoice-TTS code from this repository, stating that after release they found instances of use inconsistent with the stated intent. The TTS row in the repository model table still reads Disabled under Quick Try, while the weights remain on Hugging Face. In practice that means the TTS inference path today comes from community reimplementations and the HF weights rather than a maintained official script, and the published effort of the team has moved to ASR (released 2026-01, BitNet edge engine 2026-07, streaming variant 2026-09). Add the model card statement that it is intended for research and development only and is not recommended for commercial or real-world use without further testing, plus the biases and errors inherited from the Qwen2.5-1.5B base, and research is the honest maturity grade. The commercially deployable half is ASR, not TTS. That is the single most consequential boundary on this asset, and the headline about an open Microsoft voice model is exactly what makes it easy to miss.
The problem it solves: long audio is the real context boundary for speech models
VibeVoice is the open frontier voice model family from Microsoft, repository microsoft/VibeVoice (MIT, 54,531 stars, last push 2026-09-03, verified by us through the GitHub API). It does not present itself as another TTS model but as a family covering both synthesis and recognition: TTS for long-form multi-speaker synthesis, Realtime for streaming low latency, and ASR for structured transcription of hour long audio. All three share one technical core, which is why we carry this as one asset rather than three.
The problem is concrete: a speech model spends its context at a rate set by the frame rate. Audio has to pass through a tokenizer before an LLM can handle it, and mainstream speech tokens sit between 25 Hz and 50 Hz. Converting duration into tokens explains why long audio has been an industry gap:
| Frame rate | Tokens for 90 minutes | Tokens for 60 minutes | Fits a 64K context |
|---|---|---|---|
| 50 Hz (common discrete codebooks) | 270,000 | 180,000 | no |
| 25 Hz (the CosyVoice2/3 tier) | 135,000 | 90,000 | no |
| 7.5 Hz (VibeVoice) | 40,500 | 27,000 | yes (60 minutes single pass) |
7.5 Hz is the source of everything this family can do. The design note is that both the acoustic and the semantic tokenizers are continuous rather than quantised into discrete codebooks, which preserves audio fidelity at an ultra-low frame rate while cutting the compute cost of long sequences. With that precondition, generating 90 minutes in one pass and transcribing 60 minutes in one pass are two directions of the same fact rather than two separate engineering feats.
Three product lines: ceilings and latency
| Model | Size | Capability | Key numbers | Official entry point |
|---|---|---|---|---|
| VibeVoice-TTS-1.5B | 1.5B | Long-form multi-speaker synthesis | 90 minutes in a single pass, up to 4 speakers, cross-lingual, spontaneous singing | Weights on HF; the code was removed from the repository on 2025-09-05 and Quick Try reads Disabled |
| VibeVoice-Realtime-0.5B | 0.5B | Realtime streaming synthesis | About 300 ms first audible latency, streaming text input, roughly 10 minutes of long-form audio | HF weights plus an official Colab |
| VibeVoice-ASR-7B | 7B | Structured long-form transcription | 60 minutes in one pass, output covers who, when and what, custom hotwords, 50+ languages | HF, an official Transformers release, a Playground and finetuning code |
| VibeVoice-ASR-Streaming | - | Transcribe while audio arrives | Emits text per chunk instead of waiting for the recording to end, 10 languages, hotwords | Released 2026-09-03 with vLLM serving docs |
| VibeVoice-ASR-BitNet | - | CPU only edge inference | Heterogeneous quantisation (I8_S plus I2_S) compresses 4.62 GB to 1.58 GB, RTF below 1 on 3 or more CPU threads, no GPU | Released 2026-07-23 with the VibeASR.cpp engine and HF weights |
The output shape of the ASR line deserves its own note. The conventional pipeline for long audio is slice, recognise each slice, diarise, then align timestamps; four steps chained, global context is lost at every slice boundary and speaker indices drift across them. VibeVoice-ASR performs recognition, diarisation and timestamping jointly in one generation and emits a structured result, and it accepts user supplied hotwords and background context such as proper names and domain terms. For meeting notes, interview processing and contact centre QA that is the difference between usable and not.
The core: next-token diffusion and the base model behind it
The architecture is a next-token diffusion framework: an LLM handles textual context and dialogue flow (who is speaking, what emotion should follow), and a diffusion head produces high fidelity acoustic detail. The division of labour matters because turn consistency across a long conversation is a language problem rather than an acoustic one, and only the LLM side can keep four speakers distinct over 90 minutes; timbre detail goes to the diffusion head, which removes the pressure to raise the frame rate for fidelity. The TTS release is built on Qwen2.5-1.5B, and the risk statement says explicitly that it inherits the biases, errors and omissions of that base model.
On the record: the TTS code removal of 2025-09-05
This has to be written plainly, otherwise readers will plan a product around an open 90 minute TTS model from Microsoft. The repository news entry states that VibeVoice was released as an open research framework intended to advance collaboration in the speech synthesis community, that after release the team discovered instances of use inconsistent with the stated intent, and that because responsible use of AI is one of the guiding principles of the company the VibeVoice-TTS code was removed from the repository. The consequences:
- The weights remain on Hugging Face (
microsoft/VibeVoice-1.5B), but the Quick Try column in the repository model table reads Disabled. - No maintained official inference script lives in this repository, so the TTS path today runs through community reimplementations.
- The model card states it is intended for research and development purposes only and is not recommended for commercial or real-world applications without further testing and development.
- Every published move since then is on the ASR side: ASR released 2026-01, Transformers and Azure AI Foundry Labs integration 2026-03, the BitNet edge engine 2026-07, the streaming variant 2026-09.
We therefore file this asset at research maturity in the audio domain, and the water level judgement says in terms that the deployable half is ASR. The misuse risk is not abstract: the Risks and Limitations section names deepfakes and disinformation directly, and asks users to make sure transcripts are reliable, to check content accuracy, to avoid misleading use and to disclose AI involvement when sharing generated content.
Where it sits in an evaluation table written by a competitor
FunAudioLLM lists VibeVoice in the CosyVoice 3 README evaluation. The numbers are unfavourable to VibeVoice, so we pass them through unchanged: test-zh CER 1.16 with SS 74.4, test-en WER 3.04 with SS 68.9, and VibeVoice-Realtime-0.5B at test-en WER 2.05 with SS 63.3. In the same table CosyVoice3 base reads 1.21 and 78.0 for Chinese and 2.24 and 71.8 for English, while VoxCPM reads 0.93 and 77.2 and 1.85 and 72.9. Short-form zero-shot cloning is not its lane, and comparing it there is a misuse of the model. The dimensions it should be compared on are how long a single generation can be, whether four speakers stay distinct across it, and whether an hour of meeting audio can be handled in one pass. In the open tier there is close to no competition on those three.
Boundaries and failure modes
Five of them. No official TTS inference path, so the dependency is on community reimplementations and upgrades or security fixes will not arrive on their own. Research only language, which means commercial use requires separate compliance work and testing, and the repository licence must not be read as covering it: MIT covers the code, not the model card restrictions. Short-form similarity is not competitive, so for voice cloning look at CosyVoice 3 or Eleven v4 first. Inherited base model bias, since errors and omissions from Qwen2.5-1.5B carry through, and over long audio that surfaces as drift in wording and emotion in certain spans. Long-form failure is global: a 90 minute single pass has none of the partial recovery that slicing offers, where one bad segment can be regenerated, so one collapsed run costs the whole thing, and the same applies on the ASR side where a 60 minute single pass has higher memory and time cost than a sliced pipeline. The BitNet CPU path (RTF below 1 on 3 or more threads) is currently the only official answer that pushes it down to edge hardware.