SeeQ: a generalist language-conditioned value function nearly doubles long-horizon manipulation success

A CMU team (Saksham Singh et al., arXiv:2609.22085) presents SeeQ (Subtask-elicited Q-functions). Generalist robot policies remain brittle on complex long-horizon tasks with multiple stages or repeated attempts, while learning Q-functions from sparse task-level rewards entails long credit-assignment horizons, difficult Bellman backups and broad data-coverage requirements. SeeQ instead learns Q-values for the currently active subtask, shortening the value-prediction horizon and making temporal-difference learning effective; during training, subtask-level annotations already present in offline robot data provide the decomposition; at test time no human annotations or modular subtask predictor are needed - the Q-function architecture autoregressively predicts the active subtask in natural language before estimating its value. Built on a vision-language backbone, pretrained on diverse open-source manipulation data and finetuned downstream. Across four real-world tasks on two bimanual platforms, SeeQ substantially improves best-of-N policy steering (task-level TD stayed near the 9/24 base rate while SeeQ rose to 17/24), with both the subtask formulation and the best-of-N backup needed for the gain.





