Fun-CosyVoice3-0.5B: open TTS that turns controllability into an interface
cosyvoice-3
The third generation LLM based speech synthesis system from the FunAudioLLM group at Alibaba, 0.5B parameters, released as Fun-CosyVoice3-0.5B-2512 with base and RL weight sets plus training and inference scripts; the code repository has moved to QwenAudio/CosyVoice (Apache-2.0, 23,794 stars, verified through the GitHub API on 2026-09-29). We list it as the open source top row for speech in the audio domain, and the reason is not that it sounds most human but that it gives each of the four failure modes of production TTS an explicit interface: misread polyphonic characters go through pronunciation correction with Chinese pinyin or English CMU phonemes written straight into the input, wrong readings of numbers and symbols go through the built in text normalisation, collapsed long sentences go through RAS repetition aware sampling, and low latency goes through bidirectional streaming with an official first packet at 150ms, which is the right order of magnitude to sit inside a realtime conversational agent. Three readings matter on the metric sheet. The test-zh speaker similarity of 78.0 is the highest among 0.5B open models and above the human baseline of 75.5, yet still below the closed source Seed-TTS at 79.6, so open source first holds while world first does not. The test-hard column is the real strength, base 6.71 and RL 5.44 being the lowest in the table and better than the closed source 7.59, and hard is exactly long sentences and tongue twisters, so that lead maps to usable real scripts. English needs a discount: base test-en similarity 71.8 sits below the human baseline 73.4 and only the RL tier WER of 1.68 recovers it. Base and RL are two post training weights of one architecture, with all three error rates falling (zh CER down 33 percent, hard CER down 19 percent) and all three similarities slipping slightly, which is a good trade for broadcast and support workloads where one misread character is an incident, while voice fidelity work for a specific IP should audition the RL tier first. The repository also ships GRPO training scripts and a triton plus TensorRT-LLM runtime claimed at four times the speed of HF transformers, so post training is a path others can continue rather than a one off delivery, and the language surface covers nine languages with more than eighteen Chinese dialect accents, a tier the closed APIs still barely offer. Boundaries: the licence is two layered, Apache-2.0 for code while the weight terms live on the model pages and must be checked separately before commercial use; every reading comes from the evaluation of the authors themselves with no third party blind listening table to cross check and we did not recompute; zero shot cloning similarity comes from standard sets while real deployments depend on the noise and style of the reference recording. Graded C (vendor claim).
- CONFIDENCE
- Vendor Claim
- Official model card or keynote only, no independent re-test
- KEY METRIC
- test-zh 说话人相似度(0.5B 开源档,人类基线 75.5)
- Vendor Claim · 2025-12
- MATURITY
- Product
- research → demo → product → production
Our takeWe grade it C (vendor claim). The evaluation table in the README is the authors own, the test sets (test-zh, test-en, test-hard) and the CV3-Eval benchmark are both published by FunAudioLLM, and we have not recomputed a single number or run a blind listening test, so section 13.5 of our contract keeps it out of grade A. The table is still far more transparent than the usual vendor post: it lists the human baseline alongside the models (test-zh CER 1.26 and SS 75.5, test-en WER 2.14 and SS 73.4), and it puts closed competitors (Seed-TTS, MiniMax-Speech) and eleven open ones in the same grid with both an error column and a speaker similarity column for all three sets. A self evaluation with a control group can at least be argued with by the reader.
It earns a top row in the audio domain on one combination: the highest speaker similarity in the open tier at 0.5B parameters, together with the lowest error rate in the whole table on the hardest set. test-zh similarity of 78.0 sits above the human baseline of 75.5 and above every open model in the table, including VoxCPM at 0.5B (77.2), Index-TTS2 at 1.5B (76.5), GLM-TTS RL at 1.5B (76.4) and FireRedTTS2 at 1.5B (73.2), while being the smallest of them. The more consequential number is test-hard: CER 6.71 for the base model and 5.44 for the RL model are the lowest in the grid, ahead of Index-TTS2 at 7.12, CosyVoice2 at 6.83, VoxCPM at 8.87 and even closed Seed-TTS at 7.59. Tongue twisters, long clauses and rare characters are where production TTS actually fails, so leading there with a 0.5B model matters more than two extra points of similarity on the easy set.
The second reason is that it ships the post-training recipe, not only the weights. The base to RL delta is informative: test-zh CER 1.21 to 0.81, test-en WER 2.24 to 1.68, test-hard CER 6.71 to 5.44, all three error rates down, at the cost of a small similarity regression (78.0 to 77.4, 71.8 to 69.5, 75.8 to 75.0). That is a priced trade, a little timbre fidelity exchanged for intelligibility, and intelligibility is the dimension that generates the most complaints once a Chinese TTS voice goes live. The repository also carries GRPO training scripts (GRPO support for CosyVoice2 was contributed by Yuekai Zhang of NVIDIA), so this is a post-training path other teams can continue rather than a single released checkpoint. We file it under audio rather than llm, but half of its value is a training method.
Four boundaries to read before choosing it. First, English is not its strength: test-en similarity of 71.8 is below the human baseline of 73.4 and below VoxCPM at 72.9, and WER 2.24 trails VoxCPM at 1.85, so English first products need listening tests before commit. Second, the weight licence is not the code licence: the repository (now under the QwenAudio organisation, 23,794 stars, Apache-2.0, verified by us through the GitHub API on 2026-09-29) is Apache-2.0, while the weights are published on ModelScope and Hugging Face, and the model card terms must be checked separately before commercial use. Third, there is no third party blind arena result: unlike the music side of this domain there is no leaderboard to cross-check against, and both similarity and CER come from one self-run pipeline. Fourth, a 0.5B model does not come with an unlimited voice library: zero-shot cloning quality tracks the reference recording, and stability across the rarer entries in the 18+ dialect list has to be measured locally. The real moat is not the score but deterministic control: pinyin and CMU phoneme level pronunciation inpainting (polyphones can finally be pinned down), built-in text normalisation so numbers and symbols do not need a separate frontend, 150 ms bi-directional streaming, and instruct control over emotion, speed and volume. Closed APIs still do not offer all four.
The problem it solves: turning TTS from sampling luck into deliverable engineering
Fun-CosyVoice3-0.5B (release tag Fun-CosyVoice3-0.5B-2512) is the third generation LLM style speech synthesis system from the FunAudioLLM team at Alibaba Tongyi, 0.5B parameters, weights on ModelScope and Hugging Face, code repository now under QwenAudio/CosyVoice (Apache-2.0, 23,794 stars, verified by us through the GitHub API on 2026-09-29). The team frames it as zero-shot multilingual synthesis for in the wild speech, and the three generations form one continuous line: v1 turned speech into a discrete sequence an LLM can predict using supervised semantic tokens, v2 made it streaming (25 Hz frame rate, bi-directional), and the v3 title states the change directly, scaling up plus post-training (arXiv 2505.17589). This release ships base and RL checkpoints, training and inference scripts, and a Gradio space on ModelScope.
What separates it from the other TTS rows we carry is not sound quality but the control surface. In production, TTS failures almost never come from insufficient naturalness. They come from four predictable classes of error: polyphones read wrong, numbers and symbols read wrong, long sentences falling apart, and latency that cannot be pushed down when streaming is required. CosyVoice 3 gives each class an explicit interface:
- Pronunciation inpainting: Chinese pinyin and English CMU phonemes can be written directly into the input, so the reading of a polyphone is pinned by declaration instead of inferred. That converts the hardest Chinese TTS bug class from a probability problem into configuration.
- Built-in text normalisation: numbers, special symbols and assorted text formats no longer need a separate traditional frontend module. The optional
ttsfrdpackage improves normalisation quality, with WeTextProcessing as the fallback when it is absent. - Bi-directional streaming: both text-in and audio-out stream, with published first packet latency as low as 150 ms, which is the tier that plugs straight into a realtime conversational agent (for comparison, Eleven v4 Turbo is around 100 ms and VibeVoice-Realtime around 300 ms).
- Instruct control: language, dialect, emotion, speed and volume are all adjustable through natural language instructions, removing the need to train a separate speaker per style.
Reading the water level: the human baseline is what makes the table useful
The published evaluation gives both an error rate (CER or WER, lower is better) and speaker similarity (SS, higher is better), and it includes a human baseline as reference. The rows relevant to our selection:
| Model | Open | Size | test-zh CER ↓ | test-zh SS ↑ | test-en WER ↓ | test-en SS ↑ | test-hard CER ↓ | test-hard SS ↑ |
|---|---|---|---|---|---|---|---|---|
| Human baseline | - | - | 1.26 | 75.5 | 2.14 | 73.4 | - | - |
| Fun-CosyVoice3-0.5B (base) | yes | 0.5B | 1.21 | 78.0 | 2.24 | 71.8 | 6.71 | 75.8 |
| Fun-CosyVoice3-0.5B (RL) | yes | 0.5B | 0.81 | 77.4 | 1.68 | 69.5 | 5.44 | 75.0 |
| VoxCPM | yes | 0.5B | 0.93 | 77.2 | 1.85 | 72.9 | 8.87 | 73.0 |
| Index-TTS2 | yes | 1.5B | 1.03 | 76.5 | 2.23 | 70.6 | 7.12 | 75.5 |
| GLM-TTS / GLM-TTS RL | yes | 1.5B | 1.03 / 0.89 | 76.1 / 76.4 | - | - | - | - |
| VibeVoice-1.5B | yes | 1.5B | 1.16 | 74.4 | 3.04 | 68.9 | - | - |
| CosyVoice2 (previous generation) | yes | 0.5B | 1.45 | 75.7 | 2.57 | 65.9 | 6.83 | 72.4 |
| F5-TTS / Spark TTS / FireRedTTS2 / HiggsAudio-v2 | yes | 0.3-3B | 1.52 / 1.2 / 1.14 / 1.50 | 74.1 / 66.0 / 73.2 / 74.0 | 2.00 / 1.98 / 1.95 / 2.44 | 64.7 / 57.3 / 66.5 / 67.7 | 8.67 / - / - / - | 71.3 / - / - / - |
| Seed-TTS (closed) | no | - | 1.12 | 79.6 | 2.25 | 76.2 | 7.59 | 77.6 |
| MiniMax-Speech (closed) | no | - | 0.83 | 78.3 | 1.65 | 69.2 | - | - |
Three readings matter. First, test-zh similarity of 78.0 is the best open number but still below closed Seed-TTS at 79.6, so beating the human baseline of 75.5 is true, leading the open tier is true, and leading the world is not. Second, test-hard is where it wins: 6.71 and 5.44 are the lowest in the grid and ahead of closed Seed-TTS at 7.59. The hard set is long clauses, tongue twisters and error prone text, which is where real scripts live, so leading that column is a stronger usability signal than leading the easy one. Third, discount the English numbers: test-en similarity of 71.8 is below the human baseline of 73.4 and below VoxCPM at the same 0.5B size (72.9), and WER 2.24 trails VoxCPM at 1.85 until the RL checkpoint brings it back to 1.68.
The RL checkpoint: trading similarity for intelligibility
Base and RL are the same architecture with different post-training. Placed side by side, the direction is consistent across all six readings:
| Metric | base | RL | Change |
|---|---|---|---|
| test-zh CER ↓ | 1.21 | 0.81 | error rate down 33% |
| test-en WER ↓ | 2.24 | 1.68 | error rate down 25% |
| test-hard CER ↓ | 6.71 | 5.44 | error rate down 19% |
| test-zh SS ↑ | 78.0 | 77.4 | similarity down 0.6 |
| test-en SS ↑ | 71.8 | 69.5 | similarity down 2.3 |
| test-hard SS ↑ | 75.8 | 75.0 | similarity down 0.8 |
All three error rates fall while all three similarity scores slip slightly. For broadcast, customer service and audiobook work where one misread word is an incident, that trade is clearly worth it. For voice cloning where timbre fidelity to a specific identity is the product, listen to the RL checkpoint first. The more important point is that the repository ships the GRPO training scripts as well (GRPO support for CosyVoice2 was contributed by Yuekai Zhang of NVIDIA, alongside a triton and TensorRT-LLM runtime), so post-training here is a path another team can continue rather than a one-off checkpoint. That is the distinction we draw against models that release weights without the recipe.
Languages and dialects: 9 languages plus 18 Chinese dialects and accents
Coverage is 9 common languages (Chinese, English, Japanese, Korean, German, Spanish, French, Italian, Russian) and more than 18 Chinese dialects and regional accents, with the team naming Guangdong, Minnan, Sichuan, Dongbei, Shaanxi, Shanxi, Shanghai, Tianjin, Shandong, Ningxia and Gansu among them, plus multilingual and cross-lingual zero-shot cloning. For Chinese products the dialect tier is a capability closed APIs still largely do not offer, and regional broadcast, dialect audiobooks and local customer service all need it. It is delivered by one model plus instruct control rather than a dozen separate checkpoints. The tone numbered spelling used in the docs, Shan3xi versus Shan1xi, is itself evidence of the design: tone is a declared input, the same idea as phoneme level pronunciation inpainting.
Engineering: 0.5B, 25 Hz, four deployment paths
The frame rate stays at the 25 Hz introduced with CosyVoice2 (25 speech tokens per second), which is the physical precondition for having both long context and low latency. Four official deployment paths cover laptop to cluster:
# 1) Local Python; example.py covers zero-shot, cross-lingual and instruct usage
python example.py
# 2) vLLM: supports vLLM 0.11.x+ (V1 engine) and 0.9.0 (legacy)
# 0.10.x is untested and older releases cannot run CosyVoice inference
pip install vllm==v0.11.0 transformers==4.57.1 numpy==1.26.4
python vllm_example.py
# 3) TensorRT-LLM plus triton; published speedup is 4x over the HF transformers path
cd runtime/triton_trtllm && docker compose up -d
# 4) grpc or fastapi serving (Dockerfile under runtime/python, --max_conc sets concurrency)
python3 server.py --port 50000 --max_conc 4 --model_dir pretrained_models/Fun-CosyVoice3-0.5B
For inference stability it uses RAS (Repetition Aware Sampling). Repetition and skipped spans on long text are the standard failure mode of LLM style TTS, and RAS is a sampling strategy aimed exactly at it; together with the 25 Hz frame rate and the KV cache plus SDPA streaming optimisations, it is the reason long scripts are a defensible use case. The ecosystem is the FunAudioLLM family: FunASR (industrial recognition, 50+ languages, diarisation, streaming), Fun-ASR-Nano (end-to-end LLM based ASR, 31 languages, hotwords, vLLM streaming), SenseVoice (fast ASR plus emotion and audio events) and FunClip, so synthesis, recognition and clipping close the loop inside one stack. For a self-hosting team that is a concrete saving, not a marketing point.
Boundaries and failure modes
Four things to settle before shipping. The licence has two layers: the repository is Apache-2.0 while the weight terms live on the ModelScope and Hugging Face model cards, so commercial use needs a separate check and the repo licence must not be read as covering the weights. The metrics are single source: CER, WER and SS all come from the authors own pipeline, CV3-Eval is also theirs, there is no third party blind arena to cross-check against, and we have not recomputed anything. English and similarity involve a real trade: the RL checkpoint buys error rate at the cost of English similarity down to 69.5, so English first work should audition the base checkpoint. Zero-shot cloning inherits the reference recording: similarity was measured on standard evaluation sets, while in production the noise floor, length and speaking style of the reference take decide output quality directly, and the rarer dialect entries need local verification. One last note that is not about content but about citation: the repository moved from the FunAudioLLM organisation to QwenAudio, which signals that this line is now inside the Qwen organisation and is a positive maintenance signal, while older links in tutorials and issues will decay.