Skip to content
←Back to the radar

SOTA · CHANNEL

Audio & Video

Video generation and speech/audio in one channel: physical and temporal coherence for text/image-to-video, TTS naturalness and voice cloning, music generation, and ASR

Audio and video share one channel because the generation side stopped being two separate pipelines long ago, and because what a reader actually wants was never a silent clip plus a separate voice track. The frontier is joint generation: picture, ambience and speech coming out of one forward pass, with lip movement matching phonemes or the whole thing does not count. That capability belongs to video and audio at the same time, and splitting it across two entrances forces the reader to assemble one conclusion out of two pages. One nav entry per reader task ("I need a usable piece of footage") is why the channel exists. The two rulers are still not merged, which is why this page draws member segments instead of one blended ladder. Video is measured on temporal coherence (an object staying identical across frames), physical plausibility (gravity, collision, occlusion) and controllability (camera move, duration, first and last frame constraints), read through VBench, EvalCrafter and arena Elo. Audio is measured on TTS naturalness, voice cloning similarity, ASR WER and music generation preference. Blending them into one ruler collapses "how far is video" and "how far is speech" into a single answer that addresses neither. Read the current numbers in three parts. On video, the top of the ladder is Gemini Omni (create anything from any input, editing footage through successive turns of conversation, which moves the unit of delivery from one gamble to one editable asset). On the text-to-video arena board we sync (2026-09-22) the leader gemini-omni-1.1-flash sits at Elo 1516 on only 1,784 votes with CI plus or minus 15, while number two from the same family is 1513 on 26,576 votes with plus or minus 9; a 3 point gap far inside both intervals is not distinguishable, so the honest reading is that the Omni family is tied for first. On image-to-video the leader is minimax-h3 (1495, 57,112 votes). Seedance 2.5 (30s of single-pass narrative), the Kling 3.0 family and open-weight Wan sit in the same domain. On audio, split speech from music: speech is Eleven v4 Turbo (median model-side inference latency around 100ms), Gemini 3.8 Flash TTS (130 languages, voice design and cloning, at most 2 speakers per request), Eleven v3 (70+ languages, 5,000 characters per request) and the open source local VoiceStudio (42,831 stars, running a full TTS/ASR/music pipeline offline); music is Suno v5 and open weight YuE2, both carrying third-party SongBench means from WildSongBench. Kling (video+image+audio) and Seedance (video+audio) qualify for both member domains through joint generation. Their appearing in both segments is deliberate rather than duplicate collection: a reader arriving by the video ruler must see them, and so must a reader arriving by the audio ruler.

RULERVideo: VBench, EvalCrafter, physical & temporal coherence. Audio: TTS naturalness, voice cloning similarity, music generation, ASR WER

The ladder

11

Text/image-to-video, physical and temporal coherence

RULERVBench, EvalCrafter, physical & temporal coherence
Video GenerationMediaTopA

Gemini Omni: moving the unit of video delivery from one attempt to a conversational editing process

Google DeepMind's video and multimodal generation model, officially create anything from any input, starting with video: it edits video through step-by-step conversation, each edit building on the last and keeping the scene coherent, raising the unit of delivery from one shot to footage you can keep revising; the other two surfaces are real-world knowledge and arbitrary reference composition. Read the level with its votes: on the Artificial Analysis arena this site syncs (2026-09-22) gemini-omni-1.1-flash is #1 for text-to-video at Elo 1516 with only 1,784 votes and a ±15 CI, while #2 is the sibling at 1513 with 26,576 votes and ±9, a 3-point gap far smaller than either CI, so statistically indistinguishable: the honest statement is that the Omni family ties for the top. For image-to-video minimax-h3 leads (1495, 57,112 votes) with Omni #2 at 1488 on 3,734 votes. Pricing: $1.50 per million input tokens, $9.00 text and $17.50 video output, 720p about $0.10 per second and a 30-second clip roughly $3. GA to paid-tier developers, none on the free tier. Closed, not self-hostable, not fine-tunable. Graded A (confirmed), our first A-grade video asset; A means the reading is credible and checkable, not undisputed first place.

1516Arena T2V Elo(1784 票)Confirmed · 2026-09
ProductGoogle DeepMindSite
Gemini Omni: moving the unit of video delivery from one attempt to a conversational editing process
Video GenerationMediaTopC

Kling 3.0 series: native audio, 4K frames, and parameterised camera language as the moat

Kuaishou Kling was among the first Chinese video lines to reach a level where it can take commercial work, and its lasting strength is not raw text-to-video but image-to-video plus camera control: lock composition and character in the image stage, then let the video model interpolate from that fixed keyframe, with first/last-frame constraints and push/pull/pan/truck/orbit/follow parameters producing predictable clips. The 3.0 series (official release-note banner: API fully available) brings three substantive jumps - audio generated natively (5s/10s/12s audio tiers), native 2K/4K for both image and video, and storyboarding plus element reference turning multi-shot coherence from an editing problem into something the model controls; the line-up is 3.0 / 3.0 Omni / 3.0 Turbo, Kling Image 3.0 and -omni, Native 4K Video and Motion Control. The arena reading, stated plainly: image-to-video kling-v3-pro is #19 (Elo 1354), behind MiniMax h3, gemini-omni, wan3.0 and seedance-2.5. We could not obtain an authoritative release date for the 3.0 series, so the metric date is recorded only at the verifiable month, 2026-09. Closed, vendor API only; not benchmarked by us, graded C.

3.0 / Omni / Turbo当前旗舰系列Vendor Claim · 2026-09
Product快手 KuaishouSite
Kling 3.0 series: native audio, 4K frames, and parameterised camera language as the moat
Video GenerationMediaTopC

Seedance 2.5: raising the unit of video delivery from one shot to a 30-second story beat

ByteDance Seed line for video generation. The 1.x tiers were silent short clips from text or image; from 2.0 the architecture is a unified multimodal audio-video joint generator taking text, image, audio and video as inputs; 2.5 raises the unit of delivery to a 30-second story beat with two further extensions, white-model control, green-screen editing, professional camera work and performance direction. Its third-party reading is the strongest part: on the Artificial Analysis arena this site syncs (2026-09-22), dreamina-seedance-2.5-720p is #4 for image-to-video (Elo 1477), with 2.0 #5 (1479) and 2.5 #7 (1474) for text-to-video. Closed, reachable only through Dreamina / Volcano Engine / BytePlus, no self-hosting or fine-tuning; SeedVideoBench-2.0 is internal and cannot be reproduced. Not benchmarked by us; graded C (vendor-stated).

30s单次生成叙事时长Vendor Claim · 2026-09
ProductByteDance SeedSite
Seedance 2.5: raising the unit of video delivery from one shot to a 30-second story beat
Video GenerationMediaTopC

Wan: the open-weights baseline on the video side, letting the whole ecosystem iterate on its own hardware

Alibaba Tongyi's Wan releases video generation weights under Apache-2.0, making it the most important public baseline in the video domain. Its value differs from a closed service: not "best-looking output today" but that it is fine-tunable, reproducible and runnable on your own inference stack - the whole ecosystem of LoRAs, ControlNet-style conditioning, quantisation and few-step distillation builds on it. For an intelligence site, an open-weights baseline means the question "what can video generation do today" finally has a reference we can verify ourselves instead of trusting vendor-selected samples. The licence and the public repository we have verified; picture quality and physical plausibility remain C-grade vendor claims.

Apache-2.0开放权重许可Vendor Claim · 2025-07
Product阿里巴巴通义 Alibaba TongyiSiteRepo
Wan: the open-weights baseline on the video side, letting the whole ecosystem iterate on its own hardware
Image GenerationDesign & FrontendTopC

FLUX: the line that moved image generation from "produce a picture" to controllable editing with open weights

Black Forest Labs' FLUX family generates images with a rectified-flow transformer and ships Apache-2.0 dev weights, making it the most important open baseline on the image side; the Kontext tier turned "edit this part of the image on instruction without disturbing the rest" into a usable capability. Our own collected record also covers the FLUX 3 launch, which the vendor positions as a world model unifying image, video and audio - a claim we log as unverified.

Apache-2.0开放权重档位(dev)Vendor Claim · 2024-08
ProductBlack Forest LabsSiteRepo
FLUX: the line that moved image generation from "produce a picture" to controllable editing with open weights

TTS naturalness, voice cloning, music generation, ASR

RULERTTS naturalness, voice cloning, music generation, ASR WER
Speech & AudioMediaTopC

Eleven v4: the generation that pushes cloning fidelity far enough to reproduce bad training data

The new generation speech synthesis model from ElevenLabs, shipping in two variants: Eleven v4 for highest quality and Eleven v4 Turbo for real time use at a median inference latency of about 100 ms, with an official footnote stating that application and network latency are excluded. The vendor positions it above v3 on output quality, voice accuracy, consistency, emotion, delivery, audio tags and language coverage, and recommends migrating. We list it as the top closed source voice row in the audio domain because this generation changed the optimisation target from sounding better to sounding more like the source: timbre, cadence, delivery and other mannerisms are reproduced together, and the vendor concedes that accuracy and personal preference are not always the same thing and that you may still prefer how a voice sounded in v3. The engineering consequence is hard, since the burden of data cleaning moves back to the user. The showcase samples are explicitly labelled raw, with no EQ, compression, normalisation, de-essing or plosive removal, and the vendor says that if you hear clipping or plosives they are very likely in the training data and the model simply captured them accurately. What is deliverable comes down to three numbers: about 100 ms on Turbo; 10,000 characters per request, roughly ten minutes of audio, against 5,000 on v3 so long form text carries twice as many segment seams; and a coverage claim of 90 plus languages against a FAQ that enumerates 87, a gap we keep visible rather than smoothing away. Four boundaries matter. Cross language accent handling is a default behaviour change and not a toggle, so generating in a language that differs from the reference produces fluent native sounding speech in the target language, and the vendor calls a switchable version a research project with no timeline; any persona that depends on a native accent carried into a second language must be tested first. The control surface narrowed to Stability and Similarity, the Style and Speed sliders are gone, SSML is unsupported, and fine grained delivery moves to audio tags which the vendor admits are not perfect yet. Continuous model evolution is stated in writing, with training continuing after launch and behaviour possibly shifting over time, so teams treating a voice as a brand asset need periodic re-testing. Voice Design voices may also be less performative on v4. It is closed and not self-hostable, data must pass through ElevenLabs, and cloning compliance sits with the user; the self-hostable counterparts are CosyVoice 3, VibeVoice and Kokoro for speech and YuE2 for music under a non-commercial weight licence. This entry and the Eleven v3 already catalogued here are two generations of the same commercial stack rather than a replacement. Graded C (vendor claim): no verifiable third party benchmark, and no blind listening test by us.

~100msv4 Turbo 中位推理延迟(模型侧,不含应用与网络)Vendor Claim · 2026-09
ProductionElevenLabsSite
Eleven v4: the generation that pushes cloning fidelity far enough to reproduce bad training data
Speech & AudioMediaTopC

Suno: the platform that pulls the second half of music production inside one product

The most complete commercial platform on the text-to-song route: a style description, your own lyrics, a hummed melody or a recorded riff all work as a starting point, a full song with vocals and instrumentation arrives in seconds, and the same product then extends it, edits sections, restyles it, extracts stems and remasters it. We list it as the top closed entry for music in the audio domain on workflow completeness rather than peak audio quality. Most generative music products cover drafting and picking a take and stop there; Suno connects the rest through stem extraction with up to 12 stems, MIDI export and Suno Studio, and Studio speaks conventional DAW semantics rather than adding another prompt box, with a multitrack timeline, take lanes, comping, manual BPM to settle tempo drift and per clip transpose and speed. Custom Models trains up to three private style variants from six or more tracks you own, and Voices, formerly Personas, generates in your own singing timbre with a verification step. The tier split matters: the free plan covers creation (generation, lyrics, Cover, crop and fade, audio upload) while stems, Add Vocals, Voices and Custom Models require Pro or Premier and Studio is Premier only on desktop web, so real cost modelling should assume Premier. Version numbers do not track quality on third party benchmarks either: on WildSongBench v5 scores 6.8721, above v6 at 6.5562 and v6 Wild at 6.4195, with v4.5 at 6.6995 and v5.5 at 6.7150, and the vendor itself asks for capability based rather than version based description. Limits: closed with no self-hosting, so unreleased melodies and lyrics must be uploaded; no stable version semantics, meaning a regeneration can sound different after a model update; commercial rights follow the tier; control granularity sits at section and style level with no editable chord track, so theory level edits require Studio multitrack re-arrangement or an open model such as YuE2 that exposes an ABC score. Graded C (vendor claim); the benchmark numbers are submitted by m-a-p and have not been recomputed by us.

6.8721WildSongBench SongBench 均分(v5,第三方测)Vendor Claim · 2026-09
ProductionSunoSite
Suno: the platform that pulls the second half of music production inside one product
Speech & AudioMediaTopA

VoiceStudio: the commercial voice stack moved onto local hardware and opened to agents by default

An open source, fully local voice workstation (debpalash/VoiceStudio, AGPL-3.0, Python plus Electron) positioned as the local alternative to ElevenLabs: voice cloning, voice design, video dubbing, dictation, transcription and audiobook creation, with a claimed 646 languages. Created on 2026-04-09 and at 42,831 stars and 5,005 forks on 2026-09-29. We list it as the top self-hosted row in the audio domain, not because any model behind it is stronger, but because the abstraction layer is useful. Open voice capability is scattered across a dozen repositories with their own weight formats, device requirements, cloning interfaces and licences, so assembling a pipeline of clone a voice, transcribe, align, dub and produce an audiobook spends most of its effort on glue. VoiceStudio collects 17 TTS and 11 ASR engines into one registry and reports per engine which devices it runs on (CUDA, MPS, CPU), whether it can clone, and its max_ref_seconds and ref_strategy, on top of a workflow that actually runs. The agent surface matters most for this site: a local API and an MCP server ship as first class product features, the README includes an installation prompt ready to paste into Claude Code, Codex or Cursor, and npx skills add installs it as a skill, which means coding agents can call dubbing and transcription as tools rather than needing a human at the interface. The reference audio policy is equally rigorous, with a 75 second ceiling while how much of a longer clip reaches the model depends on the engine and is reported per engine, an automatic or saved transcript ignored beyond the limit, and a hard [clone_ref_too_long] error when sending more than 20 seconds plus transcript text to the default OmniVoice. One measurement point needs stating plainly: the official docs/benchmarks.md contains a real harness with per stage profiling, refusal to start when memory is insufficient, and a rule that only RTF warm and peak memory on named hardware and versions are accepted with no estimated numbers allowed, but the results column reads No verified rows yet. We therefore attach no performance figure at all and the card carries only the star count we verified ourselves through the GitHub API, which is what earns the A grade. That A asserts the number is real and nothing about being better than ElevenLabs. Four boundaries: AGPL-3.0 is strong copyleft with a network clause, so embedding it in a service offered to outsiders obliges you to release your own server side source, and closed commercial use means either process isolation calling only the local API under your own legal review or assembling the pipeline from Apache-2.0 pieces such as CosyVoice 3 or Kokoro; per-model licences are not settled by the wrapper, since the Breeze-TTS-2 weights behind audio.cpp are research and non-commercial, Supertonic-3 and PocketTTS need extra permissions and GPT-SoVITS requires its own server, so check every model before commercial use; the 646 language figure is the repository own claim and real coverage depends on the engine chosen, from 25 European languages on Parakeet through 50 plus on FunASR to the wider Whisper family, unverified by us language by language; and desktop is Electron only now, with 0.5.3 the last Tauri release, so existing users must migrate. It sits alongside Eleven v4 on the ladder rather than replacing it, one representing the quality ceiling and convenience of a closed commercial stack and the other the data sovereignty and engine substitutability of local self-hosting.

42,831GitHub Stars(2026-09-29,本站经 GitHub API 核)Confirmed · 2026-09
Productdebpalash (open source)SiteRepo
VoiceStudio: the commercial voice stack moved onto local hardware and opened to agents by default
Speech & AudioMediaTopC

Eleven v3: the commercial audio stack that covers speaking, listening and singing

ElevenLabs' flagship speech synthesis model, officially its most emotive and expressive: 70+ languages, a 5,000-character request cap, dramatic delivery and natural multi-speaker dialogue, with a separate Text to Dialogue API making multi-character dialogue an endpoint, not a prompt trick. We list it as a top row in audio for the stack behind it: the only commercial one covering speaking, listening and singing, with four TTS tiers, three ASR tiers including a Medical build and two music models. The TTS tiers are mutually exclusive: v3 is most expressive but non-realtime at 5,000 characters, Flash v2.5 is lowest latency (about 75ms, 50% cheaper per character) and Multilingual v2 is most stable for long text at 29 languages. On ASR, Scribe v2 covers 90+ languages with diarisation for 32 speakers. Latency figures are model-side only, excluding network and upstream LLM time, not SLAs. Limits: you split long text and own continuity across the seams; closed and not self-hostable, a hard barrier where data sovereignty matters. No speech board among the ones we sync, so listed alongside Gemini 3.8 Flash TTS without ranking. Graded C (vendor-stated): no WER recomputation or blind listening test by us.

70+语言覆盖Vendor Claim · 2026-09
ProductElevenLabsSite
Eleven v3: the commercial audio stack that covers speaking, listening and singing
Speech & AudioMediaTopC

Gemini 3.8 Flash TTS: turning director-level voice direction into API parameters

Google's speech generation model (gemini-3.8-flash-tts), positioned for studio-grade voice fidelity, expressive acting and long-form stability, the most overlooked being the last: audiobook failure is not an ugly sentence but timbre drift by hour three. The design separates scopes: text is strictly a script while performance direction travels on another channel and is never read aloud, fixing "(sighs)" that used to be spoken; round style goes into speech_metadata.style, word-level emotion into pipe tags. Four voice sources: prebuilt, an extended library, voice design IDs (describe it, then reuse) and replication IDs; design and replication arrived only in 3.8. Specs: 130 languages with automatic language detection, WAV by default at 24 kHz, at most 2 speakers per request and only with prebuilt voices. Billing is per audio token at 25 tokens per second, $0.50 input and $9.00 output per million through 2026-12-31 then $1.00/$18.00. Limits: 24 kHz is not mastering grade; replication compliance rests with the user. No speech board among the 12 we sync, so listed alongside ElevenLabs v3 without ranking. Graded C (vendor-stated): no intelligibility or blind listening test by us.

全支持语音设计+克隆+多说话人Vendor Claim · 2026-09
ProductGoogleSite
Gemini 3.8 Flash TTS: turning director-level voice direction into API parameters
Video GenerationMediaTopC

Kling 3.0 series: native audio, 4K frames, and parameterised camera language as the moat

Kuaishou Kling was among the first Chinese video lines to reach a level where it can take commercial work, and its lasting strength is not raw text-to-video but image-to-video plus camera control: lock composition and character in the image stage, then let the video model interpolate from that fixed keyframe, with first/last-frame constraints and push/pull/pan/truck/orbit/follow parameters producing predictable clips. The 3.0 series (official release-note banner: API fully available) brings three substantive jumps - audio generated natively (5s/10s/12s audio tiers), native 2K/4K for both image and video, and storyboarding plus element reference turning multi-shot coherence from an editing problem into something the model controls; the line-up is 3.0 / 3.0 Omni / 3.0 Turbo, Kling Image 3.0 and -omni, Native 4K Video and Motion Control. The arena reading, stated plainly: image-to-video kling-v3-pro is #19 (Elo 1354), behind MiniMax h3, gemini-omni, wan3.0 and seedance-2.5. We could not obtain an authoritative release date for the 3.0 series, so the metric date is recorded only at the verifiable month, 2026-09. Closed, vendor API only; not benchmarked by us, graded C.

3.0 / Omni / Turbo当前旗舰系列Vendor Claim · 2026-09
Product快手 KuaishouSite
Kling 3.0 series: native audio, 4K frames, and parameterised camera language as the moat

Evidence

12
2026-09-27Nemotron 3 Diarization tested: 100M-param real-time speaker diarization, better than expected locallyJapanese developer @ouchi tests NVIDIA Nemotron 3 Diarization (open-sourced Sep 23): real-time streaming accuracy is quite good. The ~100M-parameter model is built on the Streaming Sortformer architecture (31-layer Transformer + Arrival-Order Speaker Cache) and does diarization in a single forward pass: up to 8 speakers, 10 ms resolution, streaming latency down to about 320 ms, OpenMDW 1.1 license, runs on 4GB GPUs. The author observes the model accumulates per-speaker voice characteristics as it listens, so offline analysis of long files is slightly more accurate than real time; the HF demo caps at 2-minute inputs which limited quality, but running locally exceeded expectations. Tops VoiceArena Diarization-Bench (DER 14.72%) and scores 9.8% on AISHELL-4 (predecessor 27.2%). A key missing piece for voice agents knowing who is talking - directly relevant to robot voice interaction.2026-09-26VoiceStudio: Open-Source Voice Workstation That Runs on Your Own MachineA hands-on look at VoiceStudio, the open-source local voice workstation: switch across 14 TTS engines, clone a voice from a short clean sample, dub video into 646 languages, and produce audiobooks, dictation and transcripts with no audio leaving the machine, no subscription and no per-character billing.2026-09-23WorldCrafter: Video World Exploration Without 3D ReconstructionTencentARC WorldCrafter skips 3D reconstruction: the video generator queries an implicit memory by camera pose, so one image or text prompt yields minutes of consistent scene exploration. Weights and code are open.2026-09-23NetEase Youdao R2T2 + T3PO: Streaming ASR and Translation, Tested LiveA hands-on review of NetEase Youdao R2T2 realtime transcription and T3PO realtime translation: both hit #1 on their Hugging Face trending boards and chain into live streaming interpretation across language switches.2026-09-23GAE: A Geometry-Native Latent Space for 3D-Consistent World GenerationTencent ARC GAE compresses 3,072-channel geometry features into a 128-channel latent that decodes into RGB, depth, camera trajectories and point clouds from one generated state, halving camera-trajectory error and cutting FVD by up to 23%.2026-09-22JEV-Speech: Same 24-Layer Encoder, 2.18x Faster InferenceJEV-Speech is an Orukeet offshoot runtime from Oruk Labs that returns a transcript plus auxiliary non-transcript outputs in less than half the time. Warm p95 request latency drops from 91.93 to 42.23 ms on an A100 at batch 1, a 2.18x speedup measured from a decoded waveform in host memory to completed outputs back on the host, with all 24 encoder layers intact. The honest part is the trade-off table: a variant that changes neither weights nor precision reaches 1.66x with every transcript unchanged across 250 recordings, while the faster candidate adds BF16 arithmetic and encoder adaptation and made fifteen more word errors than the original on a separate seven-language test. Measurements are warm batch 1 over 250 historical clips with six balanced passes.2026-09-22Reka EdgeQ: An On-Device VLM Running Natively on the Snapdragon Hexagon NPUReka EdgeQ is an optimized on-device VLM running natively on the Qualcomm Snapdragon 8 Elite Hexagon NPU: 0.73s time to first token on images, +34 points over Gemma 4 E4B on MLVU video, 6.9 mWh per inference, and a GPU left completely idle during inference.2026-09-15Odyssey-3: A Foundation World Model for Physical AgentsOdyssey-3 is an autoregressive diffusion-transformer foundation world model that controls robot arms, humanoids, cars, and drones with task-specific experiential data.2026-09-07PKU Motion-Omni End-to-End Dialogue SystemPKU launches Motion-Omni, an end-to-end system generating speech and synchronized full-body motion with RTF 0.78.2026-09-06Vivix-W1: Streaming-Native Multimodal Model for Real-Time Interactive VideoVivix Labs introduces Vivix-W1, a streaming-native multimodal model that turns generated video into a live, responsive world. Beyond camera-control navigation, W1 accepts real-time touch, text and voice prompts during playback: tap to insert objects (a cat lands on the cafe table), direct characters, or restyle the whole scene on the fly (a New York street morphs toward an Industrial Revolution look). The 110-second demo shows continuous streaming generation that reacts to input without regenerating from scratch.2026-09-04StreamTalk Streams Speech Into SMPL-X Body GesturesStreamTalk turns audio into body gestures in real time, outputs SMPL-X pose skeletons, and leaves the 3D rig to the user; the demo is available on Hugging Face Spaces.2026-09-03GWM Worlds 2 Makes Video and Audio Into Real-Time SimulationRunway GWM Worlds 2 turns high-fidelity video and audio generation into real-time interactive simulation, letting users define environments, subjects, visual style, physical rules, and ambience.

How it is used

No deep reads yet. This section is for workflows, hands-on practice and failure notes — not news rewrites.

Assets you can use now

Boundaries & failure modes

Boundaries and failure modes, ordered from the ones specific to this channel down to the ones each member domain carries on its own. One, the easiest thing to misread at channel level: the union ladder is not a ranking. It pulls assets from both member domains into one list under a single ordering rule (featured, then maturity, then newest metric date) purely so a reader can see at a glance what we have reviewed on this line. An Elo of 1516 on video and a 100ms latency on speech are not commensurable, and appearing higher does not mean stronger. For a ranking, enter a member segment, or open /sota/video and /sota/audio and read each ruler separately. Two, missing evaluation keeps the whole channel conservative. Of the 12 boards we sync, there are text-to-video and image-to-video boards and no speech board, no TTS naturalness board and no ASR WER board; on music there is only third-party academic work such as WildSongBench, narrow in coverage and inconsistent in sampling budget across vendors (the 6.9632 for YuE2 is best-of-8 while the 6.8721 for Suno v5 is not the same budget, so the two numbers cannot be compared directly). Every water-level judgement on the audio side therefore carries vendor-claim; confirmed is reserved for verifiable public facts (GitHub star counts, licence, price, self-reported latency), and we run no blind listening test and recompute no WER. Three, alignment error in joint audio-video generation accumulates. A lip-sync miss is tolerable at 5 seconds and unmistakable across a 30 second narrative, and no third party publishes a protocol for measuring it, so vendor demos are always hand-picked samples. Read every "video with sound in one pass" claim as vendor-claim. Four, duration and physics remain the hard wall on the video side. Single-pass generation generally stops at 5 to 10 seconds; anything longer is built by relaying first and last frames or by stitching segments, and style drift and character inconsistency at the seams have to be picked out by hand. Fluids, cloth, multi-person interaction and hand-to-object contact are the frequent failure points, because what the model learns is the texture statistics of looking physical, not conservation laws. Five, most boundaries on the audio side are not technical. A few seconds of sample is enough to clone a voice, while consent chains, watermarking and abuse detection are all immature; open-weight TTS can be withdrawn over abuse (there is a 2025-09 precedent), so "open and usable today" is not "available long term". We fold compliance status into the confidence judgement. On music, training-data copyright litigation is unresolved, so check the licence before any commercial delivery. Six, latency and real-time behaviour are barely measured in evaluations. First-packet latency for high-quality TTS typically runs from a few hundred milliseconds to seconds, and genuinely usable real-time conversation needs streaming synthesis plus barge-in, so the most natural model on a board may be unusable in a call. Long-text consistency (paragraph prosody, how numbers and proper nouns are read, ambiguous readings, mixed Chinese and English text) is the most common production failure, while evaluation sets usually read only short sentences. Seven, cost stacks along two lines, so do not budget one side only. 720p video runs about $0.10 per second, roughly $3 for a 30 second clip, and three to ten re-rolls is normal in commercial delivery; speech is billed per second or per character, where bulk TTS unit prices sit far below video but the cost of a real-time channel lives in concurrency rather than duration; music generation is close to the video magnitude, with a single track taking tens of seconds to a few minutes. A finished narrated piece with a score must be budgeted as video re-rolls plus multiple voice takes plus multiple music takes, not as three unit prices added together.