π0.7 is a steerable generalist robotic foundation model conditioned on diverse multimodal context—language commands, subgoal images, and task metadata—enabling strong out-of-the-box performance in unseen environments, zero-shot cross-embodiment transfer, and emergent capabilities matching RL-finetuned specialists.
This paper addresses the frame mismatch in VLA models between camera-frame observation and robot-frame action by introducing robot-centric pointmaps—images whose pixels store 3D coordinates in the robot frame. Pointmaps provide robot-frame 3D geometry while preserving image structure, enabling cross-viewpoint generalization across diverse camera setups.