Skip to content
←Back to Open Source

OPEN SOURCE DEEP DIVE

Image GenerationText-to-ImageDiffusion Transformer

SANA: 4K Text-to-Image That Fits in Laptop VRAM

SANA is NVIDIA Labs open-source codebase for high-resolution image and video generation (9k+ stars, Apache-2.0, PyTorch). Its efficiency comes from two structural changes: linear attention replaces vanilla attention in the transformer, and a deep compression autoencoder compresses images 32x instead of the usual 8x. Official figures: the 0.6B model renders a 1024 image in 0.9 s, roughly one fortieth of FLUX-dev 23 s, while FID improves from 10.15 to 5.61. The series now spans one-step generation, video, world models and streaming editing, and runs 4K in 8 GB of VRAM under 4-bit quantization. No benchmarks reproduced; recorded as unverified.

NVlabs/Sana9.2kPythonApache-2.05 min read

What it is

SANA is NVIDIA Labs' open-source codebase of efficient diffusion models for high-resolution image and video generation, with complete training and inference pipelines. The repository went public in October 2024, has over nine thousand stars, and ships under Apache 2.0 in Python and PyTorch. Its original claim was efficient high-resolution image synthesis with a linear diffusion transformer; two years on it has grown into a series, with image, video, world-model and post-training lines all living in the same repo.

On the paper side: ICLR 2025 Oral, ICML 2025, ICCV 2025 Highlight, ICLR 2026 Oral. The line worth remembering is in its own introduction: 20× smaller and 100× faster than Flux-12B, while still reaching 4K.

Where the efficiency comes from: two cuts at once

The usual way to make a diffusion transformer fast is fewer steps or fewer parameters. SANA cuts at the structural level, and the two cuts reinforce each other.

The first replaces vanilla attention in the transformer with linear attention. Standard attention costs grow with the square of sequence length, and that bill gets ugly as resolution rises; linear attention pulls the relationship back to roughly linear, saving exactly the most expensive part at high resolution.

The second cut is the autoencoder. It uses a deep compression autoencoder, DC-AE, compressing images 32× where conventional diffusion models typically use 8×. Compression ratio directly sets the number of latent tokens — fewer tokens, and the savings from the first cut get multiplied again.

Two more pieces come with it: the text encoder is a modern decoder-only LLM using in-context learning for better text-image alignment, and the sampler is a Flow-DPM-Solver to cut sampling steps. Together these are what supports the "20× smaller, 100× faster" claim.

The numbers: speed without paying for it in quality

The table below is the official comparison at 1024×1024. We have not reproduced it; treat this tier as unverified. Note the counter-intuitive part: it is not merely faster, several quality metrics improve at the same time.

ModelThroughput (img/s)Latency (s)Params (B)SpeedupFID ↓CLIP ↑GenEval ↑
FLUX-dev0.0423.012.01.0×10.1527.470.67
Sana 0.6B1.70.90.639.5×5.6128.800.68
Sana 1.6B1.01.21.623.3×5.9228.940.69
Sana 1.5 1.6B1.01.21.623.3×5.7029.120.82
Sana 1.5 4.8B0.264.24.86.5×5.9929.230.81

The 0.6B Sana produces a 1024 image in 0.9 seconds — roughly one fortieth of FLUX-dev's 23 seconds — while FID drops from 10.15 to 5.61 (lower is better) and both CLIP score and GenEval land above FLUX. That tells you the efficiency comes from structure rather than from trading quality for speed, which is by far the more common story.

Stacking more compute still pays: the 1.5-series 4.8B has the best CLIP score (29.23), yet its GenEval is edged out by the 1.6B in the same series (0.81 versus 0.82). The project positions 1.5 as training-time and inference-time compute scaling — what that tier sells is scalability, not a single optimum.

A whole series grown from images

The repository now covers far more than text-to-image:

SANA-Sprint uses continuous-time consistency distillation (sCM) to compress generation to one or a few steps, reported at 0.1 s per 1024 image on an H100 and 0.3 s on a consumer RTX 4090. SANA-Video and LongSANA use block linear attention and causal mix-FFN for long videos. SANA-Video 2.0 comes in 5B and 14B sizes with hybrid linear and softmax attention plus attention residuals. SANA-WM is a 2.6B controllable world model producing 720p, one-minute continuous scenes with six-degree-of-freedom camera control. SANA-Streaming is a 2B real-time streaming video-to-video editor. Sol-RL is the post-training method: NVFP4 rollout sampling with BF16 optimization, reported at 4.64× faster convergence.

One engineering result deserves its own mention: Sol Engine reached 3.95× acceleration on MiniMax-H3, a 33B omni-modal audio+video transformer, on GB200 — and more on desk-side hardware (3.92× on DGX Spark, 4.52× on an RTX 5090) — reportedly in 4.5 hours of optimization, with no distillation, no LoRA and no calibration pass.

Deployment limits: 4K and 8 GB of VRAM

The deployment figures are specific: the 1.6B 4K model produces a 4096×4096 image within 20 seconds; DC-AE tiling brings 4K inference inside 22 GB of GPU memory; and with 8-bit or 4-bit quantization plus model offload, 4K runs in 8 GB. The stated overall claim is deployable on laptop GPUs with under 8 GB of VRAM via 4-bit quantization.

Ecosystem support is in place too: merged into Hugging Face diffusers (SanaPipeline from 0.32.0), ComfyUI nodes, SGLang integration, ControlNet, LoRA and DreamBooth training, 4-bit and 8-bit quantization, and NVIDIA's own Cosmos-RL for post-training.

How to read the numbers

In the official video table, SANA-Video 2.0 5B's 13.2 s corresponds to 480×832×81 frames at 40 steps on a single H100 under the VBench protocol; the same model's end-to-end latency at 736×1280×81 is 30.9 s. The 4-step variant is a DMD preview, not final. The 14B config and checkpoint are explicitly not included yet. The repository states all of this itself — just don't compare numbers across different protocols side by side.

Also in that table, SANA-Video-2B scores 84.05 on VBench with 2B parameters in 36 seconds, above Wan-2.1-14B's 83.73 (1897 s, 14B). Figures from the project; we have not reproduced them.

Unverified

We ran no benchmarks and no reproduction. Every figure above comes from the repository and its documentation site. Stars, release cadence — public for two years, still shipping new versions — and first-class support in diffusers make it a layer worth cataloguing on the image-generation line, but its capability tier is recorded as unverified.

As an Amazon Associate, we earn from qualifying purchases.