
VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
VLA-Precision from USTC tackles the two bottlenecks of real-world online RL for large VLAs — value-signal-induced policy drift and large-model compute overhead. The ACoB algorithm establishes asymmetric co-bootstrapping across timescales: early intervention-guided BC lifts performance fast, global return propagation and local preference ranking progressively calibrate value estimates, and relative-advantage improvement with reference regularization suppresses drift — with local ranking specifically fixing the overestimation of overwritten human-corrected proposals. ACoB-Stream delivers up to 10.9x throughput via invariant-state decoupling and on-demand streaming. Across 9 high-precision chemistry tasks on 4 embodiments: 98.3% mean success in 45.8 min/task, 27.6 s episodes (1.2x and 1.8x VLA/RL baseline speed).