Skip to content
← Tags

#Audio Tags (1)

Speech & AudioMediaTopC

Eleven v4: the generation that pushes cloning fidelity far enough to reproduce bad training data

The new generation speech synthesis model from ElevenLabs, shipping in two variants: Eleven v4 for highest quality and Eleven v4 Turbo for real time use at a median inference latency of about 100 ms, with an official footnote stating that application and network latency are excluded. The vendor positions it above v3 on output quality, voice accuracy, consistency, emotion, delivery, audio tags and language coverage, and recommends migrating. We list it as the top closed source voice row in the audio domain because this generation changed the optimisation target from sounding better to sounding more like the source: timbre, cadence, delivery and other mannerisms are reproduced together, and the vendor concedes that accuracy and personal preference are not always the same thing and that you may still prefer how a voice sounded in v3. The engineering consequence is hard, since the burden of data cleaning moves back to the user. The showcase samples are explicitly labelled raw, with no EQ, compression, normalisation, de-essing or plosive removal, and the vendor says that if you hear clipping or plosives they are very likely in the training data and the model simply captured them accurately. What is deliverable comes down to three numbers: about 100 ms on Turbo; 10,000 characters per request, roughly ten minutes of audio, against 5,000 on v3 so long form text carries twice as many segment seams; and a coverage claim of 90 plus languages against a FAQ that enumerates 87, a gap we keep visible rather than smoothing away. Four boundaries matter. Cross language accent handling is a default behaviour change and not a toggle, so generating in a language that differs from the reference produces fluent native sounding speech in the target language, and the vendor calls a switchable version a research project with no timeline; any persona that depends on a native accent carried into a second language must be tested first. The control surface narrowed to Stability and Similarity, the Style and Speed sliders are gone, SSML is unsupported, and fine grained delivery moves to audio tags which the vendor admits are not perfect yet. Continuous model evolution is stated in writing, with training continuing after launch and behaviour possibly shifting over time, so teams treating a voice as a brand asset need periodic re-testing. Voice Design voices may also be less performative on v4. It is closed and not self-hostable, data must pass through ElevenLabs, and cloning compliance sits with the user; the self-hostable counterparts are CosyVoice 3, VibeVoice and Kokoro for speech and YuE2 for music under a non-commercial weight licence. This entry and the Eleven v3 already catalogued here are two generations of the same commercial stack rather than a replacement. Graded C (vendor claim): no verifiable third party benchmark, and no blind listening test by us.

~100msv4 Turbo 中位推理延迟(模型侧,不含应用与网络)Vendor Claim · 2026-09
ProductionElevenLabsSite
Eleven v4: the generation that pushes cloning fidelity far enough to reproduce bad training data