Eleven v3: the commercial audio stack that covers speaking, listening and singing
elevenlabs-v3
ElevenLabs' flagship speech synthesis model, officially its most emotive and expressive: 70+ languages, a 5,000-character request cap, dramatic delivery and natural multi-speaker dialogue, with a separate Text to Dialogue API making multi-character dialogue an endpoint, not a prompt trick. We list it as a top row in audio for the stack behind it: the only commercial one covering speaking, listening and singing, with four TTS tiers, three ASR tiers including a Medical build and two music models. The TTS tiers are mutually exclusive: v3 is most expressive but non-realtime at 5,000 characters, Flash v2.5 is lowest latency (about 75ms, 50% cheaper per character) and Multilingual v2 is most stable for long text at 29 languages. On ASR, Scribe v2 covers 90+ languages with diarisation for 32 speakers. Latency figures are model-side only, excluding network and upstream LLM time, not SLAs. Limits: you split long text and own continuity across the seams; closed and not self-hostable, a hard barrier where data sovereignty matters. No speech board among the ones we sync, so listed alongside Gemini 3.8 Flash TTS without ranking. Graded C (vendor-stated): no WER recomputation or blind listening test by us.
- CONFIDENCE
- Vendor Claim
- Official model card or keynote only, no independent re-test
- KEY METRIC
- 语言覆盖
- Vendor Claim · 2026-09
- MATURITY
- Product
- research → demo → product → production
Our takeWe grade it C (vendor claim). The grade is constrained by the evidence form available across the whole audio domain: what we can verify is specification and pricing (language counts, character limits, latency tiers, feature lists), while "most expressive", "state-of-the-art speech recognition" and "35% fewer errors on clinical audio" are vendor statements. We have not measured WER, not run blind preference listening, and not measured prosodic continuity after long-form segmentation, so we do not upgrade them to A. With no third-party board in this domain at all, C is the only honest grade.
What earns it a place at the top of the ladder is not one model but stack completeness: four TTS tiers, three ASR tiers and two music tiers make it the only commercial vendor covering speaking, listening and music in one place. Real audio production was never only synthesis - a dubbed video needs TTS, subtitles need ASR word-level timestamps and diarisation, and a score needs a music model that can inpaint locally. Having all three behind one API and one account system saves integration cost rather than unit price. The Scribe v2 parameters deserve particular attention: 1000-term keyterm prompting, 65 entity types and 32-speaker diarisation is a configuration aimed squarely at enterprise meetings and professional transcription, not at a general demo.
Selection has to respect the trade-off in that table: expressiveness, latency and language coverage are mutually exclusive and no tier wins all three. v3 is the most expressive yet capped at 5,000 characters; the most stable long-form model, Multilingual v2, covers 29 languages; the ~75ms Flash v2.5 is narrower on both other axes. The right approach is one model per scenario, not one "best" model everywhere. Two risks must stay visible. First, every latency figure is model-side and excludes network and upstream LLM time, and in a real voice agent the LLM segment usually dominates, so they are not end-to-end SLAs. Second, it is closed with no self-hosting: audio must pass through ElevenLabs, which is a hard gate for data-sovereignty-sensitive work in medical, government and financial recording even with a Scribe v2 Medical tier. Compliance for voice cloning always stays with the operator - capability is not authorisation.
The problem it solves: audio as a complete production line rather than a TTS endpoint
Eleven v3 is ElevenLabs' flagship speech synthesis model, described officially as "our most emotionally rich, expressive speech synthesis model", with four properties: dramatic delivery and performance, 70+ languages, a 5,000 character limit per request, and support for natural multi-speaker dialogue.
We catalogue it at the top of the audio ladder not for that model alone but for what sits behind it: the only commercial audio stack that covers speaking, listening and music in one place - TTS (v3 / v3 Conversational / Multilingual v2 / Flash v2.5), ASR (Scribe v2 / Realtime / Medical) and music (Eleven Music v2.5 / v2). Competition in speech stopped being about whose timbre sounds most human; it is about who can put dubbing, subtitling and scoring into one API and one account system.
Tiering: expressiveness, latency and stability pull against each other
| Model | Latency | Languages | Chars per request | Use |
|---|---|---|---|---|
| Eleven v3 | Non-realtime tier | 70+ | 5,000 (about 5 minutes of audio) | Audiobooks, podcasts, dubbing: when you need performance |
| Eleven v3 Conversational | ~280ms | 70+ | - | Realtime dialogue agents, with audio tags for fine-grained control |
| Eleven Multilingual v2 | - | 29 | 10,000 | The most stable tier for long-form generation |
| Eleven Flash v2.5 | ~75ms | 32 | 40,000 | Telephony and low-latency work; 50% lower price per character on the API |
The most useful thing in that table is the trade-off it exposes: the most expressive model, v3, allows only 5,000 characters per request; the most stable long-form model, Multilingual v2, covers 29 languages; the lowest-latency Flash v2.5 is narrower on both expressiveness and language coverage. No tier wins all three axes, so the engineering answer is to pick per scenario rather than choose one "best" model and use it everywhere. The official guidance points the same way: v3 for multi-character audio experiences, long-form narration and emotional dialogue, shipped alongside a dedicated Text to Dialogue API that treats multi-speaker dialogue as an interface rather than a prompting trick.
Read the ~280ms and ~75ms figures with their scope in mind: both are model-side latencies as published (carrying footnotes in the docs), excluding network round trip and excluding the time your upstream LLM takes to emit text. In a real voice agent the LLM segment usually dominates.
The ASR side: Scribe v2's parameter density is worth more attention than the TTS
- Accurate transcription in 90+ languages, with smart language detection.
- Keyterm prompting, up to 1000 terms: feed in a domain vocabulary and proper nouns and terminology hold up. This is a hard requirement for enterprise audio (medical, legal, internal codenames) and exactly where generic ASR fails.
- Entity detection across 65 entity types: the output is typed text, not just text.
- Word-level timestamps plus speaker diarisation for up to 32 speakers: that number is clearly aimed at meetings and multi-party interviews.
- Scribe v2 Realtime: ~150ms streaming transcription, same 90+ languages, word-level timestamps and entity detection.
- Scribe v2 Medical: fine-tuned for clinical audio, with 35% fewer transcription errors on clinical audio than Scribe v2, equal accuracy on everyday speech, and identical features, languages, pricing and API. This is one of the few commercial audio products we have seen that quantifies what vertical fine-tuning buys.
The music side: Eleven Music v2.5
v2.5 is the current top tier, described as better quality and prompt adherence with richer melodies, deeper arrangements and more layered instruments. It keeps the same workflows as v2 and therefore the three capabilities that matter in production: composition plans, audio reference, and inpainting. v2 is positioned as studio-grade music from natural-language prompts in any style, with complete control over genre, style and structure.
Inpainting is the easiest of the three to underestimate and the biggest time saver: if four bars of a cue are wrong, regenerating the whole piece gambles the parts you already liked. Being able to edit locally is what makes it a production tool.
Boundaries
- v3 caps at 5,000 characters per request, about five minutes of audio. Long form means segmenting it yourself, and prosodic continuity and timbre consistency across segment boundaries are your problem; the model does not guarantee cross-request consistency.
- The three axes are mutually exclusive: expressiveness, latency and language coverage - see the table. There is no all-round tier.
- Closed source, no self-hosting: no downloadable weights, no deployment in your own environment, and audio must pass through ElevenLabs. For data-sovereignty-sensitive work (medical, government, financial recordings) that is a hard gate even with a vertical Scribe v2 Medical tier.
- Latency figures are model-side: ~280ms / ~150ms / ~75ms all carry official footnotes and exclude network and upstream LLM time, so they are not end-to-end SLAs.
- "SOTA" and "35% fewer errors" are vendor statements: we did not recompute WER on our own test sets and ran no blind comparison against Gemini TTS, OpenAI TTS or the Whisper family.
- Compliance for voice cloning sits with the operator: capability is not authorisation, and requirements for voice rights and synthetic-audio labelling differ by jurisdiction.
- No third-party speech board: none of the twelve leaderboards this site syncs covers speech or audio, so this asset cannot be ranked.
Our verification status
Facts come from the ElevenLabs official Models documentation, read in full: the four TTS tiers (v3 at 70+ languages / 5,000 characters / multi-speaker dialogue; v3 Conversational at ~280ms with audio tags; Multilingual v2 at 29 languages / 10,000 characters; Flash v2.5 at ~75ms / 32 languages / 40,000 characters / 50% lower price per character), v3's stated use cases and the Text to Dialogue API, the three ASR tiers (Scribe v2 at 90+ languages / 1000 keyterms / 65 entity types / word-level timestamps / 32-speaker diarisation; Realtime at ~150ms; Medical at 35% fewer clinical errors), and the two music tiers (v2.5 with composition plans, audio reference and inpainting; v2 with genre and structure control).
Confidence is graded C (vendor claim), for the same reason as Gemini 3.8 Flash TTS: everything checkable here is specification and pricing, not a quality reading. "Most expressive", "state-of-the-art speech recognition" and "35% fewer errors on clinical audio" are vendor statements. We have not measured WER, not run blind preference listening, and not measured prosodic continuity after segmenting at the 5,000-character boundary. Reaching grade A would require WER and speaker-similarity measurement on a fixed multilingual test set, organised blind comparison, and a quantified measure of consistency decay over segmented long form.
We place it alongside Gemini 3.8 Flash TTS at the top of the audio ladder rather than ordering the two. Their strengths genuinely differ: ElevenLabs wins on stack completeness (commercial products across speaking, listening and music, with an especially dense ASR parameter set) and on latency tiering; Gemini 3.8 Flash TTS wins on voice design plus cloning plus inline emotion control inside one model, with broader language coverage (130 vs 70+) and published, lower audio-output pricing. With no third-party board in this domain, we state facts side by side and do not invent a ranking.