SoL-Refiner: one-step refinement turns low-res video into 4K - NVIDIA speed-of-light video refiner

NVIDIA Research (Efficient AI Team and Singapore Lab; arXiv:2609.37969; Haozhe Liu et al.) presents SoL-Refiner. High-resolution video generation is costly because inference scales with the number of spatiotemporal tokens; generating low-resolution first and refining is the practical alternative, but conventional multi-step refinement introduces a second sampling bottleneck. SoL-Refiner transforms low-resolution outputs from diverse base generators into 4K video with a single denoising step. Its three-stage recipe combines high-resolution continual training, frame-based RL post-training, and one-step distribution-matching distillation with both bidirectional and streaming inference. The authors also introduce Refiner-Bench, a video refinement benchmark built from the outputs of different video generators, with a shared-input protocol comparing refiners at roughly 2K output resolution. At 2K, the one-step SoL-Refiner outperforms all external refiners on VBench and UniPercept averages; at 3840x2176 it improves both metrics over the three-step LTX-2.3 Refiner; the complete acceleration stack delivers an 8.91x speedup in refinement latency over the same baseline in the 2K latency setting. Measured on the project page: the MiniMax H3 deployment pipeline drops from 152.3s to 5.64s (27.03x; five-second 1344x768 at 24 fps on GB200 GPUs), and the SANA-Video compute pipeline from 20.04s to 12.54s (1.60x; 1x H100, 81 frames at 16 fps).





