Flex-pi: A Multi-Stream World-Action Model with Compute Flexibility

A 6B-parameter manipulation policy from the University of Washington and the Allen Institute for AI. Instead of predicting only future RGB image latents, it jointly denoises three streams in a shared latent space: RGB appearance, 3D geometry as pointmaps, and object-centric DINO semantics. A single 5B trunk plus a 1B action expert sits on frozen encoders (a Wan-2.2 video VAE for RGB and pointmaps, frozen DINOv3), and per-stream dropout with cross-modality forcing lets one checkpoint run 56 input/output combinations, from an action-only fast path to full joint world-and-action generation. Evaluated on real bimanual YAM-workcell tasks (plate racking, utensil sorting, kitchen organization, gripper self-repair, soft-bag zipping) plus RoboTwin and LIBERO, against pi-0.5, ManiFlow, Fast-WAM and RDT. The argument: pixel-reconstruction latents carry no explicit signal for the 3D geometry or object semantics that manipulation needs. This is the JEPA critique, predict in a semantic latent space rather than reconstruct pixels, applied to robot action.





