Caltech ROM-Nav: Decoupling Navigation from Embodiment Lets Humanoids Walk Better

Caltech and Amazon Safe Autonomy Frontiers Lab present ROM-Nav (Learning Safe Humanoid Navigation from Reduced Order Models), a two-stage navigation system for humanoids. Stage one trains a navigation policy on a drastically simplified body — a single integrator with heading, essentially a moving dot colliding with an occupancy grid — learning spatial reasoning, exploration and multi-floor route-finding without ever contending with balance, foot placement or stairs. Stage two distills that into the humanoid via weighted PPO with a KL-divergence loss against the reduced-order policy. The system perceives through a Mid-360 LiDAR and ZED Mini depth camera via frozen pretrained denoising-VAE CNN encoders feeding self- and cross-attention, a GRU and an MLP. It issues velocity commands at 5 Hz on top of a frozen locomotion controller at 50 Hz, with every command passing through a Poisson safety filter that synthesizes a control barrier function from the live point cloud and projects via closed-form QP. Runs on a Unitree G1, trained in 12 hours (reduced-order) plus 32 hours (full) on a single H100.





