EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data

Egocentric human data is abundant, but human motion is not always positive supervision for robot policies because of embodiment gaps: naive behavior-cloning co-training can hurt performance. From Georgia Institute of Technology, CoRL 2026. EgoWAM is a controlled human-robot co-training framework that fixes the policy backbone, action head and data mixture while varying only the world prediction target, comparing Pixel, DINO and 3D motion flow. The key finding is that the state-prediction branch of a World Action Model bridges the embodiment gap, letting robot performance scale with diverse in-the-wild human data. Built on a Heterogeneous Pretrained Transformer backbone with embodiment-specific stems feeding a shared trunk, plus a conditional flow matching action head and a swappable world-model head. Across three real-world bimanual tasks, pixel-based prediction transfers weakly while DINO improves out-of-distribution object and scene generalization by up to 4x and 3D flow improves in-domain performance by 20-30%. Under deliberately unaligned human data, behavior cloning drops below its robot-only baseline while 3D-flow world-model co-training stays robust.





