0:48EgoEngine turns egocentric human videos into robot demos@Randle_Liu · 113 views · 2026-06-22EgoEngineEgocentric VideoZero-Shot
CReF: Cross-modal and Recurrent Fusion for Depth-conditioned Humanoid LocomotionCReF is a single-stage depth-conditioned humanoid locomotion framework that maps onboard proprioception and forward-facing depth directly to joint position targets, without explicit geometric intermediates. It couples proprioception and depth tokens via proprioception-queried cross-modal attention, fuses them with a gated residual block, and integrates temporal context with a GRU regulated by a highway-style output gate. A terrain-aware foothold placement reward extracts supportable candidates from foot-end point-cloud windows and rewards touchdowns near them. Full CReF leads all terrain categories and difficulty levels in simulation and transfers zero-shot to a physical AGIBOT X2 Ultra across handrail stairs, hollow pallets, reflective interference, and cluttered outdoor scenes.Yuan Hao, Ruiqi Yu, Shixin Luo·Mar 31, 2026HumanoidLocomotiondepth-conditionedMar 31, 2026