FreeVideo: MiniMax H3 video generation on a laptop with as little as 8GB VRAM — VDN-H3 linear attention plus hardware-adaptive execution planning, open source as a ComfyUI plugin

Haocheng Xi (UC Berkeley) released FreeVideo (FlashML-org, Apache-2.0): a local inference engine for MiniMax H3 on consumer GPUs, running the 33B omni-modal video model with as little as 8GB VRAM and 16GB RAM. It builds on OpenVDN's 8-step VDN-H3 model — a hybrid-attention Video DeltaNet rework of H3 that replaces quadratic long-range attention with a linear Video Delta Attention branch. The core is hardware-adaptive execution planning: VRAM, system memory, and disk are scheduled as a single memory hierarchy; the engine picks native FP8 or FP8-storage-with-BF16-compute per GPU architecture and auto-probes available attention kernels (SageAttention, etc.); weight streaming, asynchronous prefetching, and chunked computation keep peak memory at 8GB. Shipped as a ComfyUI plugin: the Windows launcher installs and starts ComfyUI with a dedicated creative workspace supporting two-pass sampling, batch generation, and history; power users can switch to node view to add LoRAs or customize workflows; Linux also offers a CLI (./freevideo generate). Multimodal inputs cover text prompts, first/last frames, and image/video/audio references. Weights carry the MiniMax H3 Community License (territorial and acceptable-use restrictions); the code is open source





