Eleven v3: the commercial audio stack that covers speaking, listening and singing
ElevenLabs' flagship speech synthesis model, officially its most emotive and expressive: 70+ languages, a 5,000-character request cap, dramatic delivery and natural multi-speaker dialogue, with a separate Text to Dialogue API making multi-character dialogue an endpoint, not a prompt trick. We list it as a top row in audio for the stack behind it: the only commercial one covering speaking, listening and singing, with four TTS tiers, three ASR tiers including a Medical build and two music models. The TTS tiers are mutually exclusive: v3 is most expressive but non-realtime at 5,000 characters, Flash v2.5 is lowest latency (about 75ms, 50% cheaper per character) and Multilingual v2 is most stable for long text at 29 languages. On ASR, Scribe v2 covers 90+ languages with diarisation for 32 speakers. Latency figures are model-side only, excluding network and upstream LLM time, not SLAs. Limits: you split long text and own continuity across the seams; closed and not self-hostable, a hard barrier where data sovereignty matters. No speech board among the ones we sync, so listed alongside Gemini 3.8 Flash TTS without ranking. Graded C (vendor-stated): no WER recomputation or blind listening test by us.