0:43Flex-pi: A Multi-Stream World-Action Model with Compute Flexibility@LeoKharon · 115 views · 2026-09-08VLAWorld ModelsRobot Manipulation
See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action ModelsThis paper addresses the frame mismatch in VLA models between camera-frame observation and robot-frame action by introducing robot-centric pointmaps—images whose pixels store 3D coordinates in the robot frame. Pointmaps provide robot-frame 3D geometry while preserving image structure, enabling cross-viewpoint generalization across diverse camera setups.Byungkun Lee, Dongyoon Hwang, Dongjin Kim·Jul 13, 2026VLAPointmapcross-viewpointJul 13, 2026