PAPER DEEP DIVE
TurboT2VA: Score-Regularized Consistency Distillation for Fast Large-Scale Text-to-Video-Audio Generation
Joint text-to-video-audio generation is prohibitively expensive. TurboT2VA distills a 19B joint video-audio model via per-modality normalization and a progressive curriculum (discrete consistency warm-up, continuous refinement, joint consistency-distribution matching). On LTX-2, four-step distillation cuts 512x768 latency from 50.52s to 2.51s (20.1x) while preserving quality, audio fidelity, diversity and AV sync; an architecture-aware stack (guarded W8A8, fused ops, padded-text compaction, modality-aware sparse attention) reaches 5.83s from 318.74s at 1024x1792 on one H20 (54.67x).
1. Background: the inference cost of joint video-audio generation
Joint text-to-video-audio (T2VA) generation (e.g. LTX-2) produces synchronized visuals and sound in one pass, but relies on multi-step diffusion/flow sampling; a 19B joint Transformer (14B video backbone + 5B audio backbone) takes tens of seconds per sample. Existing acceleration routes each have weaknesses: dCM (discrete-time consistency) builds targets over a finite timestep grid, suffering discretization errors and grid-scheduling sensitivity; sCM (continuous-time consistency) takes the continuous limit and follows the local tangent of the teacher trajectory with more faithful supervision, but its JVP-based tangent estimation is numerically fragile at 19B scale; DMD (distribution matching distillation) improves realism at the cost of diversity. Combining them for large-scale T2VA remained unexplored.
2. Joint cross-modal distillation: not two summed modality losses
TurboT2VA formulates distillation over paired video-audio latents: a shared base timestep $t$, independently sampled noises, TrigFlow noising: $$z_t^m = \cos(t)\, z_0^m + \sin(t)\, \epsilon^m,\quad m \in \{v,a\}$$ The student predicts both modalities in one forward: $(\hat z_{0,\theta}^v, \hat z_{0,\theta}^a) = G_\theta(z_t^v, z_t^a, t, c)$. Key design: losses are computed per modality, but the forward graph is shared — one cross-modal Transformer receives paired latents with shared conditioning and timestep embeddings, so video and audio distillation gradients flow through the same model and cross-modal alignment is optimized implicitly.
The sCM stage requires a joint JVP of the student field whose direction passes through both video and audio tokens: $$J_\theta^m = c_t\Big(\frac{\partial F_\theta^m}{\partial t} + \frac{\partial F_\theta^m}{\partial z_t^v}F_T^v + \frac{\partial F_\theta^m}{\partial z_t^a}F_T^a\Big)$$ The video direction depends on audio tokens and vice versa — the essential difference from simply adding two independent modality losses. The sCM direction is per-modality L2-normalized (Eq. 6) to prevent video gradients from dominating.
3. Joint DMD and modality-balanced optimization
After the student learns a stable joint trajectory, DMD is introduced: on paired on-policy student samples $\hat z_0=(\hat z_v^0,\hat z_a^0)$, perturb to $\hat z_t$ and take the discrepancy between the fake-score network and the teacher prediction as the score difference $g_{\mathrm{DMD}}^m$, using a residual-scale normalizer (Eq. 11-12) rather than sCM's L2 one. Because the DMD term is computed from paired samples generated by the same student under the same condition and timestep, it improves realism without breaking the cross-modal structure learned by joint sCM. The final objective: $$\mathcal{L}_{\text{stage3}} = \lambda_{\text{sCM}} \mathcal{L}_{\text{sCM}}^{\text{joint}} + \lambda_{\text{DMD}} \mathcal{L}_{\text{DMD}}^{\text{joint}}$$ Both objectives reuse the same batch, text condition and paired latents; balancing is applied only at the loss level — the student remains a single unified generator.
4. Three-stage progressive curriculum: dCM → sCM → sCM+DMD
Direct sCM+DMD from the base initialization is trainable but suboptimal: early DMD pushes the student to match teacher-like samples before it has built a sufficiently broad trajectory prior, weakening inherited diversity. The curriculum:
| Stage | Objective | Role |
|---|---|---|
| Stage 1: dCM warm-up (2K steps) | Discrete consistency | No tangent estimation, numerically stable; builds a coarse denoising prior and stable timestep conditioning |
| Stage 2: sCM refinement (0.5K steps) | Continuous teacher trajectory | Removes discretization error, transfers teacher trajectory directions faithfully, preserves diversity |
| Stage 3: sCM+DMD joint (4.5K steps) | Trajectory + distribution | sCM maintains trajectory structure and diversity while DMD anchors the teacher distribution for perceptual quality |
Training totals 7K steps: 8×H20 GPUs, 100K text-video-audio samples at 512×768×121 frames, AdamW (lr $2\times10^{-5}$, $\beta_1=0$, $\beta_2=0.999$), BF16 + FSDP + gradient checkpointing, CFG scales 3.0 (video) / 5.0 (audio), about 21 hours. LTX-2 originally uses rectified-flow velocity parameterization; a TrigFlow wrapper (Eq. 16) maps latents and timesteps without touching the backbone.
5. Architecture-aware inference stack: compressing per-step cost beyond distillation
Few-step distillation reduces the number of Transformer evaluations, not the cost of each; at high resolution per-step cost dominates. A video-only stack is unsafe on the joint model (four attention paths differ in sequence lengths, masks and matrix shapes), so the authors design:
| Component | Mechanism | Notes |
|---|---|---|
| Modality-aware sparse attention | SageSLA only on video/audio self-attention (query-dependent block map, retention $\rho=0.3$, floor $\tilde\rho=\max(\rho, 1/N_{blk})$ guarantees one block) | Bidirectional cross-modal and masked text cross-attention stay dense, preserving AV exchange and text conditioning |
| Shape-aware W8A8 linears | Static per-output-channel weight + dynamic per-row activation quantization; TileLang kernel accumulates INT32 and applies both scales + bias in the epilogue | Concatenated QKV projections where inputs are shared; unsupported dtype/shape falls back to BF16 |
| Fused multimodal kernels | Modulated RMSNorm, adaptive scale-shift, gated residual, split RoPE fusion | Every kernel guards dtype/contiguity/shape/mask and falls back to the original path |
| Padded-text compaction | At batch 1 the shared text mask compacts video/audio conditioning to valid tokens | Removes only padded positions; disabled for batch > 1 |
6. Experiments
Standard resolution (512×768, 200 prompts): the 4-step student runs 2.51s vs the 40-step teacher's 50.52s (20.1×), Javis 0.1963 vs 0.195, Visual 2.389 vs 1.968, Desync 0.388 vs 0.617 — several AV metrics beat the teacher. Competitive against open-source JavisDiT (3.1B) / OVI (11B) / DaVinci-MagiHuman (15B) and closed-source Kling v3 / Sora 2 / Veo 3.
| Metric | Teacher (40 steps) | TurboT2VA (4 steps) |
|---|---|---|
| Latency (s) ↓ | 50.52 | 2.51 (20.1×) |
| Javis ↑ / Desync ↓ | 0.195 / 0.617 | 0.1963 / 0.388 |
| VBench aesthetic / imaging | 0.5286 / 0.5884 | 0.5519 / 0.6520 |
| TTA-Bench production quality | 6.577 | 6.584 |
Diversity ablation (8 seeds per prompt, ImageBind embeddings, 5,600 pairs per modality): sCM-only is the most diverse (V-A Avg Div 0.4032) but Javis only 0.1131; DMD-only reaches 0.1812 Javis with the lowest diversity (0.2691); staged sCM+DMD gives the best balance (Javis 0.1963 / Div 0.3259). Qualitatively: sCM-only changes subjects and layouts but loses the requested cat–robot interaction; DMD-only produces polished yet repetitive compositions; the staged model keeps both.
High-resolution deployment (1024×1792, 121 frames, one H20, batch 1, generator-only): dense 40-step teacher 318.74s → +W8A8+fused ops 233.34s (1.37×) → swap to the 4-step student 12.16s (26.2×) → +SageSLA(ρ=0.3)+padded-text compaction 5.83s (54.67×). The unoptimized 4-step student takes 16.50s, so the architecture-aware stack contributes an extra 2.83×. Sweeping ρ from 0.5 to 0.2 moves latency 6.44→5.57s with Javis essentially flat (0.1939–0.1948); ρ=0.3 gives the best CAVP and lower desync.
| Config (cumulative) | Latency (s) ↓ | Speedup |
|---|---|---|
| Dense 40-step teacher | 318.74 | 1.00× |
| + W8A8 + fused ops | 233.34 | 1.37× |
| + 4-step TurboT2VA student | 12.16 | 26.20× |
| + SageSLA(ρ=0.3) + padded-text compaction | 5.83 | 54.67× |
Training-curve ablations: the staged curriculum overtakes direct sCM+DMD once the joint stage begins and keeps the Javis advantage through 7,500 global steps; from the same dCM+sCM prefix, DMD refinement improves quickly then plateaus below the joint objective, while sCM-only stays stable but clearly lower — staged sCM+DMD peaks at 0.1963 at 7K steps.
7. Conclusion
TurboT2VA pushes score-regularized continuous-time consistency distillation to 19B-scale joint video-audio generation for the first time: modality-balanced losses solve video domination, the dCM→sCM→sCM+DMD curriculum structures "learn the trajectory first, refine quality later", and the architecture-aware stack compresses per-step cost. Model-level (few-step) and system-level (per-step) improvements are orthogonal — 20.1× from distillation, 54.67× from the full stack. Code and demos: thu-ml/TurboDiffusion.