BLOG
Isaac 0.5 is Perceptron's open-source embodied foundation model with 36B sparse parameters: it reads images, video, language, robot state and previous actions to answer video questions, point and track objects, report task progress, and generate robot actions. The team establishes a scaling law trading video for teleop: scaling general video from 1,000 to 1M hours cuts the teleoperation needed for action loss 2.50 from ~5,900 hours to 28 (210x). Trained on 35+ robot systems, 100K hours of robot experience, 1M hours of video and 3T multimodal tokens, it introduces semantic world modeling (predicting future percepts), the mHarmony typed multimodal interface, and Null Experts for dynamic compute — leading all five perception task families at 8.5x lower inference cost. Weights, training code and LeRobot inference code are fully released.
1X shares 1XWM progress: action-controllable video generation predicts future outcomes of robot actions, shifting policy evaluation from physical experiments to simulated forecasting at scale.
1X introduces 1XWM, a video-pretrained world model integrated into NEO as a robot policy. Unlike VLAs, 1XWM derives robot actions from text-conditioned video generation, leveraging world dynamics from internet video for zero-shot generalization without large-scale teleoperation data.
GRASP proposes three innovations making long-horizon planning with world models practical: lifting states for parallel-in-time optimization, stopping brittle state gradients while keeping action gradients, and periodic sync refinement. Achieves leading success rates at long horizons on Push-T.
DYNA Robotics introduces Dyna-2, a world-action model pre-trained on 1M+ hours of egocentric human video, demonstrating the first human-to-robot transfer scaling law. With just hours of fine-tuning, it performs tasks across embodiments, achieving 87% zero-shot pass rate at customer sites.
Riemann Dynamics releases Riemann-1.0: a fully causal autoregressive World Action Model that unifies executable robot policy and action-conditioned world simulation in one architecture. Through progressive pretraining on 232K+ hours of heterogeneous embodied experience, it achieves SOTA on LIBERO (99.0%), RoboTwin 2.0 (94.3%), RoboCasa365 (62.6%), and real-world manipulation (85.0% SR).