
WARP-RM: Warp-Augmented Relative Progress Reward for Data Curation
WARP learns dense signed relative progress from successful demos via time-warp augmentations, then WARP-BC filters and reweights action chunks for behavior cloning. On bimanual T-shirt folding, throughput rises up to ~18× vs vanilla BC as suboptimal demos increase; in 512 paired sim bottle scenes it reaches 290 bottles/hr.
Justin Yu, Andrew Goldberg, Kavish KondapJun 26, 2026
Imitation LearningData CurationBehavior CloningJun 26, 2026