Self-Adaptive VLA: learning from its own failed rollouts to fix itself

Hardware shifts are everywhere in real deployment - wear, manufacturing tolerances, imperfect calibration. Under actuation bias or joint encoder offsets, the base VLA policy drops from about 100% to about 0% success even though nothing about the task changed. The post-training recipe: 1) roll out the frozen base policy under injected hardware shifts; 2) pair each rollout with expert demos pre-compensated for that same shift (no new human data); 3) train a lightweight plug-in encoder that compresses a rollout into a single context token. The token modulates the DiT action head through AdaLN, summed with timestep conditioning; computed once per trial, so zero extra inference cost during closed-loop control. Context tokens can be summed across trials - each failure reveals a shift the previous one masked: the first fixes the grasp, the second the insertion, the third succeeds. Across 4 precision bimanual/dexterous tasks (Piper, Piper-X, Marvin arms, 20-DoF Wuji hand): one trial of context recovers 46-49% of lost performance, ensembling up to 6 trials recovers 80-84%; on cross-workstation deployment where the base policy fails 5/5, two failed trials as context yield 5/5 success.





