Skip to content
Tags
Speech & AudioMediaTopC

Eleven v3: the commercial audio stack that covers speaking, listening and singing

ElevenLabs' flagship speech synthesis model, officially its most emotive and expressive: 70+ languages, a 5,000-character request cap, dramatic delivery and natural multi-speaker dialogue, with a separate Text to Dialogue API making multi-character dialogue an endpoint, not a prompt trick. We list it as a top row in audio for the stack behind it: the only commercial one covering speaking, listening and singing, with four TTS tiers, three ASR tiers including a Medical build and two music models. The TTS tiers are mutually exclusive: v3 is most expressive but non-realtime at 5,000 characters, Flash v2.5 is lowest latency (about 75ms, 50% cheaper per character) and Multilingual v2 is most stable for long text at 29 languages. On ASR, Scribe v2 covers 90+ languages with diarisation for 32 speakers. Latency figures are model-side only, excluding network and upstream LLM time, not SLAs. Limits: you split long text and own continuity across the seams; closed and not self-hostable, a hard barrier where data sovereignty matters. No speech board among the ones we sync, so listed alongside Gemini 3.8 Flash TTS without ranking. Graded C (vendor-stated): no WER recomputation or blind listening test by us.

70+语言覆盖Vendor Claim · 2026-09
ProductElevenLabsSite
Eleven v3: the commercial audio stack that covers speaking, listening and singing
Speech & AudioMediaTopC

Gemini 3.8 Flash TTS: turning director-level voice direction into API parameters

Google's speech generation model (gemini-3.8-flash-tts), positioned for studio-grade voice fidelity, expressive acting and long-form stability, the most overlooked being the last: audiobook failure is not an ugly sentence but timbre drift by hour three. The design separates scopes: text is strictly a script while performance direction travels on another channel and is never read aloud, fixing "(sighs)" that used to be spoken; round style goes into speech_metadata.style, word-level emotion into pipe tags. Four voice sources: prebuilt, an extended library, voice design IDs (describe it, then reuse) and replication IDs; design and replication arrived only in 3.8. Specs: 130 languages with automatic language detection, WAV by default at 24 kHz, at most 2 speakers per request and only with prebuilt voices. Billing is per audio token at 25 tokens per second, $0.50 input and $9.00 output per million through 2026-12-31 then $1.00/$18.00. Limits: 24 kHz is not mastering grade; replication compliance rests with the user. No speech board among the 12 we sync, so listed alongside ElevenLabs v3 without ranking. Graded C (vendor-stated): no intelligibility or blind listening test by us.

全支持语音设计+克隆+多说话人Vendor Claim · 2026-09
ProductGoogleSite
Gemini 3.8 Flash TTS: turning director-level voice direction into API parameters