
SparkDiffusion: Mitigating the High-Sparsity Trap — A Unified Framework for up to 265× Single-GPU Acceleration of Visual Generation
SparkDiffusion (arXiv:2609.23153) from Peking University, Tsinghua and Alibaba diagnoses the high-sparsity trap: at 97% attention sparsity, step-local training loss keeps falling while terminal video quality stagnates or degrades — the root cause is supervision, since flow-matching-style step-local losses cannot constrain terminal errors (an oracle probe shows fixing the five highest-noise steps removes most terminal error). The three-stage recipe: a short sparse warm-up (compensated sparse attention, RoLA), few-step trajectory-mixed distillation (CrossDistill: high-noise PCM consistency + low-noise DMD distribution matching, 3-step CFG-free student), and FP8 W8A8 quantization with fused kernels. It sustains 97% sparsity on long 720P generation, achieving a measured 265x end-to-end speedup on Wan2.1-T2V-14B-720P on a single RTX 5090 (220x on H100) and 1.3 s for 1.3B-480P videos, with seed-level diversity closest to the dense reference (VBench 83.15 vs 83.69). Code open-sourced.