
Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning
Q-Planning freezes a large behaviour-cloning policy and trains an approximately one-billion-parameter off-policy Q function. At inference, Q performs single-step weighted averaging over BC action-chunk draws; during online self-improvement, both successful and failed rollouts update only Q. Ten iterations raise mean LIBERO success from 92.1% to 97.6%, while two bimanual real-robot tasks improve from 40%/25% to 90%/80%.
Varun Giridhar, Anant Khandelwal, Jeremy A. CollinsAug 21, 2026
Q-LearningVLARobot ManipulationAug 21, 2026
