Light-O1: Whole-Body Intelligence Scaled by Human Action Pretraining

Light Origins releases Light-O1, its first general-purpose embodied foundation model. Structured human actions recovered from internet video are tokenized into a unified humanoid action representation (root trajectory, body pose, hand state), interleaved with language and visual observations, and pretrained autoregressively on a 4B base at budgets from 3.75B to 120B multimodal tokens, about 100k hours of human action. The result is a cross-embodiment transfer scaling law: next-action-token loss and whole-body pose error fall as power laws after adaptation to Nymeria egocentric data, Unitree G1 teleoperation (HIW-500) and in-house LightBot data. Adapted through decoded unified actions or a diffusion action expert plus RLHF, the open Light-O1-Preview checkpoint (Qwen3_5ActionForConditionalGeneration, Apache-2.0) outputs (frames, 138) actions at 20 FPS, scores 79.3% macro success on 24 RoboCasa GR-1 tabletop tasks against pi0.5 70.9%, DIAL 70.2% and GR00T N1.7 59.1%, and reaches Elo 1472.8 in a 30k-prompt motion arena.





