
BeyondRetarget: Learning Executable Humanoid Motions Directly from Monocular Video
BeyondRetarget removes the SMPL intermediate representation and maps monocular RGB video end-to-end to executable humanoid motions: 18 semantic keypoints with per-link non-uniform scaling form unified supervision, a shared representation decodes onto 8 humanoids, and contact-aware refinement yields 26.4mm error, zero collapse, and 193ms latency enabling real-time visual teleoperation.
Tianyu Xiong, Yi Lu, Jinrui WangSep 24, 2026
HumanoidMotion retargetingImitation LearningSep 24, 2026