Eleven v4: the generation that pushes cloning fidelity far enough to reproduce bad training data
elevenlabs-v4
The new generation speech synthesis model from ElevenLabs, shipping in two variants: Eleven v4 for highest quality and Eleven v4 Turbo for real time use at a median inference latency of about 100 ms, with an official footnote stating that application and network latency are excluded. The vendor positions it above v3 on output quality, voice accuracy, consistency, emotion, delivery, audio tags and language coverage, and recommends migrating. We list it as the top closed source voice row in the audio domain because this generation changed the optimisation target from sounding better to sounding more like the source: timbre, cadence, delivery and other mannerisms are reproduced together, and the vendor concedes that accuracy and personal preference are not always the same thing and that you may still prefer how a voice sounded in v3. The engineering consequence is hard, since the burden of data cleaning moves back to the user. The showcase samples are explicitly labelled raw, with no EQ, compression, normalisation, de-essing or plosive removal, and the vendor says that if you hear clipping or plosives they are very likely in the training data and the model simply captured them accurately. What is deliverable comes down to three numbers: about 100 ms on Turbo; 10,000 characters per request, roughly ten minutes of audio, against 5,000 on v3 so long form text carries twice as many segment seams; and a coverage claim of 90 plus languages against a FAQ that enumerates 87, a gap we keep visible rather than smoothing away. Four boundaries matter. Cross language accent handling is a default behaviour change and not a toggle, so generating in a language that differs from the reference produces fluent native sounding speech in the target language, and the vendor calls a switchable version a research project with no timeline; any persona that depends on a native accent carried into a second language must be tested first. The control surface narrowed to Stability and Similarity, the Style and Speed sliders are gone, SSML is unsupported, and fine grained delivery moves to audio tags which the vendor admits are not perfect yet. Continuous model evolution is stated in writing, with training continuing after launch and behaviour possibly shifting over time, so teams treating a voice as a brand asset need periodic re-testing. Voice Design voices may also be less performative on v4. It is closed and not self-hostable, data must pass through ElevenLabs, and cloning compliance sits with the user; the self-hostable counterparts are CosyVoice 3, VibeVoice and Kokoro for speech and YuE2 for music under a non-commercial weight licence. This entry and the Eleven v3 already catalogued here are two generations of the same commercial stack rather than a replacement. Graded C (vendor claim): no verifiable third party benchmark, and no blind listening test by us.
- CONFIDENCE
- Vendor Claim
- Official model card or keynote only, no independent re-test
- KEY METRIC
- v4 Turbo 中位推理延迟(模型侧,不含应用与网络)
- Vendor Claim · 2026-09
- MATURITY
- Production
- research → demo → product → production
Our takeWe grade it C (vendor claim). ElevenLabs publishes no verifiable third party benchmark: there is no MOS figure, no speaker similarity number and no WER we can cite, only the vendor own generational comparison plus a set of showcase samples, and we have run no blind listening test of our own. Contract section 13.5 therefore forbids an upgrade to A. What makes this entry credible within that grade is something else: the vendor labels its showcase samples as raw, with no EQ, compression, normalisation, de-essing or plosive removal, and states outright that if you hear clipping, plosives, uneven EQ or volume jumps then those artefacts are very likely in the training data and v4 simply captured them accurately. Publishing the flaws of your own output for anyone to hear is unusual in an audio model release, and it is why we list this as the top closed source voice row rather than merely relaying the marketing claim.
The case for the top of the ladder is a generational step on the cloning fidelity axis, paired with a real time latency that is rare at this quality tier. The change from v3 to v4 is not that it sounds better but that it sounds more like the source: timbre, cadence, delivery and other mannerisms are reproduced together, and the vendor itself concedes that accuracy and personal preference are not always the same thing and that you may still prefer how a voice sounded in v3. That concession amounts to admitting the optimisation target of this generation is fidelity rather than pleasing the ear. The engineering consequence is hard: the burden of data cleaning moves back to the user, because reference recording quality now determines output quality directly instead of being masked by model side beautification, and the vendor changed its training data guidance accordingly (stick to a single speaking style, with an explicit note that the guidance may change). What is actually deliverable comes down to three numbers: median inference latency around 100 ms on the Turbo variant, with an official footnote stating that application and network latency are excluded; 10,000 characters per request, roughly ten minutes of audio; and a language claim of 90 plus against a FAQ that enumerates 87. For contrast, v3 allows 5,000 characters per request so long form text carries twice as many segment seams, and Flash v2.5 is faster at about 75 ms but covers only 32 languages with text normalisation off by default. We keep the gap between 90 plus and 87 visible rather than smoothing it, because a headline coverage claim and an enumerable list have always been two different numbers.
Four boundaries must be understood before migrating. Cross language accent handling is a default behaviour change, not a toggle: when the generated language differs from the reference voice language, v4 produces fluent native sounding speech in the target language instead of carrying the reference accent across, and the vendor says making it switchable is still a research project with no timeline. If a product persona depends precisely on a speaker carrying a native accent into a second language, as in character voiceover or interview reconstruction, this upgrade silently changes the output and must be tested or avoided by staying on v3. The control surface narrowed: only Stability and Similarity remain, the Style and Speed sliders are gone, SSML is unsupported, and fine grained delivery moves to inline audio tags which the vendor admits are not perfect yet and remain an active area of investment. For a pipeline built on SSML or on the Style slider this is migration cost, not upgrade benefit. Continuous model evolution is stated in writing: training continues after launch, behaviour may shift over time and periodic re-testing is recommended, so any product that treats a voice as a brand asset should put that on an operations calendar because reproducibility of output is not under your control. In addition Voice Design voices may be less performative on v4 in the vendor own words, so projects built on Voice Design should compare both generations before moving.
Then the one that cannot be routed around: closed and not self-hostable, data must pass through ElevenLabs, and the compliance burden for voice cloning sits with the user because capability is not authorisation. For teams whose data cannot leave their own infrastructure, or who need to hold the voice weights themselves, the self-hostable counterparts in our audio domain are CosyVoice 3, VibeVoice and Kokoro for speech, and YuE2 for music with the caveat that its weights are CC-BY-NC-4.0 and therefore non-commercial. One selection note: v4 and the Eleven v3 already catalogued on this site are two generations of the same commercial stack, not a replacement. The vendor recommends migrating while also conceding that a few edge cases fit v3 better, so two coexisting rows on the ladder is the correct reading.
The problem it solves: pushing clone fidelity far enough that it reproduces the flaws in your training data
Eleven v4 (eleven_v4) and Eleven v4 Turbo (eleven_v4_turbo) are the new generation of ElevenLabs speech synthesis models. The vendor states that v4 improves output quality, voice accuracy, consistency, emotion, delivery, audio tags and language coverage over v3, and recommends migrating. It sits on the same commercial audio stack as the Eleven v3 entry already in this catalogue rather than replacing it: v3 allows 5,000 characters per request while v4 allows 10,000, about ten minutes of audio, which halves the number of segmentation seams in long form work.
The most informative change in this generation is the direction of cloning fidelity. The documentation is unusually blunt: v4 reproduces the EQ, loudness, plosives, sibilance and volume fluctuation present in the reference recording, and the showcase samples are raw with no post-processing at all, meaning no EQ, compression, normalisation, de-essing or plosive removal. The vendor warns that recurring issues heard in those samples are most likely present in the training data for that voice and the model simply captured them accurately. The engineering consequence is hard: v4 pushes data cleaning back onto the integrator, because reference recording quality now determines output quality directly instead of being masked by model smoothing. The accompanying guidance changed as well, and currently recommends a single speaking style in training audio, with the vendor stating it is still experimenting and that this guidance may change.
Cross-language accent handling: a deliberate behaviour change that must be tested before migrating
This is the largest behavioural difference from every older model. When the generated language matches the language of the reference voice, the original accent is preserved exactly as before. When the two differ, v4 produces fluent natural speech in the target language rather than carrying the reference accent across. Cloning a Korean speaker and generating English yields natural English rather than English spoken with a Korean accent. That makes a single voice usable across every supported language and is the most practical gain in this generation for dubbing and multilingual voice agents.
It is, however, a default rather than an option. The vendor says it is exploring turning it into a toggle, describes that as a research project and gives no timeline. Products whose persona depends on a voice carrying its native accent into another language, such as character dubbing or interview reconstruction, therefore have to test v4 directly before migrating and may need to stay on v3. This site calls the change out separately because it belongs to the class of upgrades that silently alter output rather than purely improving it.
Two variants and a narrower control surface
| Model | Positioning | Latency (model side) | Chars per request | Languages |
|---|---|---|---|---|
| Eleven v4 | Highest quality: content creation, audiobooks, character voiceover | Non-realtime tier | 10,000, about ten minutes | Vendor states 90+, the FAQ enumerates 87 |
| Eleven v4 Turbo | Realtime: conversational agents, interactive voice | Median ~100ms | - | Same family coverage |
| Eleven Flash v2.5 (previous realtime tier) | Lowest latency and lowest price | ~75ms | 40,000 | 32 |
| Eleven v3 Conversational | Previous expressive realtime tier | ~280ms | - | 70+ |
The control surface narrowed in this generation. Only two settings remain, Stability, which controls how consistent delivery stays across generations, and Similarity, which controls adherence to the reference voice at some cost to naturalness. The Style and Speed sliders are not available and SSML is not supported. In their place are audio tags such as [whispering], [shouting] and [laughing], which the vendor says v4 handles with nuance beyond previous models while also stating that they are not perfect yet and remain an active area of investment. This is a paradigm shift from continuous sliders to inline markers, and for a pipeline built on SSML or the Style slider it is a migration cost rather than an upgrade benefit.
Both the ~100ms and ~75ms figures carry the official footnote: they exclude application and network latency. In a real voice agent the upstream LLM usually dominates end to end latency by a wide margin, so these numbers are not SLAs. Flash v2.5 also disables text normalisation by default to protect latency, with an Enterprise-only apply_text_normalization flag, so the correct pattern on a low latency tier is to have the upstream LLM normalise numbers and symbols first.
Limits: the model keeps moving, Voice Design got weaker, closed and not self-hostable
- The vendor states explicitly that the model will keep evolving. Training continues after launch and behaviour may shift over time, with a recommendation to re-test your use case periodically. For products where a voice is a brand asset this belongs in the operations calendar, because output reproducibility is not under your control.
- Voices created with Voice Design may be less performative on v4 than on earlier models, in the words of the documentation. Optimising for faithful cloning of a real voice is not the same as optimising for designing a voice from a description, so Voice Design projects should compare both generations before migrating.
- Two cloning paths. Instant Voice Cloning takes a one to two minute sample and is usable within seconds with no separate training step; Professional Voice Cloning trains a dedicated model on longer recordings and is fully supported on v4, with existing PVCs fine-tunable for v4 from Voice Lab.
- Closed and not self-hostable, so all audio passes through ElevenLabs and the compliance burden for voice cloning stays with the integrator, since capability is not authorisation. The self-hostable counterparts in this domain are CosyVoice 3, VibeVoice and Kokoro for speech and YuE2 for music under a non-commercial licence.