PAPER DEEP DIVE
VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
VLA-Precision from USTC tackles the two bottlenecks of real-world online RL for large VLAs — value-signal-induced policy drift and large-model compute overhead. The ACoB algorithm establishes asymmetric co-bootstrapping across timescales: early intervention-guided BC lifts performance fast, global return propagation and local preference ranking progressively calibrate value estimates, and relative-advantage improvement with reference regularization suppresses drift — with local ranking specifically fixing the overestimation of overwritten human-corrected proposals. ACoB-Stream delivers up to 10.9x throughput via invariant-state decoupling and on-demand streaming. Across 9 high-precision chemistry tasks on 4 embodiments: 98.3% mean success in 45.8 min/task, 27.6 s episodes (1.2x and 1.8x VLA/RL baseline speed).
TL;DR
Pretrained VLA models enable broad manipulation but remain unreliable on tasks demanding precision and repeatability. Real-world online RL can push VLAs past the demonstration ceiling, but hits two bottlenecks: unreliable value signals induce policy drift, and large-VLA overhead constrains throughput and sample efficiency. VLA-Precision tackles both with the ACoB algorithm and the ACoB-Stream architecture. ACoB establishes asymmetric co-bootstrapping across timescales: early intervention-guided behavioral learning rapidly improves the policy while enhancing online experience quality; as autonomous experience accumulates, global return propagation and local preference ranking progressively calibrate value estimates, yielding relative action advantages for reference-regularized policy improvement while suppressing drift. ACoB-Stream organizes around invariant-state decoupling and on-demand streaming, delivering up to 10.9× throughput and compute-efficiency gains. Across nine high-precision chemistry tasks in four categories and four robot embodiments: 98.3% mean success in 45.8 min per task, with 27.6 s episodes at 1.2× and 1.8× the speed of VLA and RL baselines.
Figure 1: Framework overview. VLA-Precision combines rapid behavioral learning, progressive value calibration, and reference-regularized relative-advantage policy improvement within a closed-loop experience–policy architecture.
Background: Why Real-World Online RL Is Hard
Recent VLA models have acquired broadly transferable capabilities through large-scale pretraining on heterogeneous multimodal data. With only a few task-specific demonstrations they perform diverse manipulation and generalize across environmental variations. But broad adaptability cannot guarantee precise and repeatable task completion: small residual errors at critical stages still cause intermittent failures.
Behavior cloning learns task behavior directly from dense action labels, but errors compound outside the demonstrated distribution and performance remains bounded by demonstration quality. RL instead uses temporal credit assignment to learn long-term action values from interaction, so applying RL to VLA post-training enables continual optimization through interaction-derived return signals, mitigating compounding error and surpassing the demonstration ceiling.
Existing real-world RL methods split by model scale: RL for large pretrained VLAs and RL for compact robot policies. The paper identifies two critical bottlenecks that persist: (1) algorithmic stability — unreliable value estimates induce policy drift; and (2) system efficiency — large-VLA computational overhead limits training throughput and reduces overall online RL efficiency.
Prior mitigations regularize RL updates with BC constraints from demonstrations or interventions. While these preserve prior behaviors and stabilize exploration, they provide limited improvement in value estimation. ConRFT's Q maximization sacrifices prior performance and permits drift; Robo-Dopamine's large progress-based reward model can be misled by OOD hallucinations; RL-100 needs 12.1 hours of real-robot rollout per task on average; π*0.6 exceeds 90% success on most tasks but requires over 1,000 real-robot rollouts per task; RL Token accelerates critical-phase throughput up to 3× but still needs 400–1,000 online episodes per task, with externalized improvements and policy handoffs limiting end-to-end adaptation.
Method: ACoB Algorithm + ACoB-Stream Architecture
Problem Formulation: Two-Stage Post-Training
Figure 2: Two-stage pipeline. Stage I performs full-parameter imitation learning on demonstrations to obtain the task-specific prior; Stage II initializes real-world online RL from it under an asynchronous actor–learner process.
Physical interaction under task instruction $\ell$ is modeled as a standard MDP:
$$\mathcal{M}_\ell = (\mathcal{S}, \mathcal{A}, P, r, \rho_0, \bar{\gamma})$$
At decision step $t$, the state $s_t = (o_t, q_t, \ell) \in \mathcal{S}$ comprises visual observation $o_t$, robot state $q_t$, and instruction $\ell$. A VLA action is an $H$-step action chunk $a_t = (u_{t,0},\dots,u_{t,H-1})\in\mathcal{A}\subseteq\mathbb{R}^{H\times d}$. Executing it induces $s_{t+1}\sim P(\cdot\mid s_t,a_t)$ and yields:
$$R^{(H)}_t = \sum_{h=0}^{H-1} \gamma^h r_{t,h}$$
with primitive-step discount $\gamma$ and $\bar{\gamma}=\gamma^H$. The policy is initialized from task-finetuned $\pi_{0.5}$, and only the LoRA parameters $\theta\in\mathbb{R}^p$ in the action expert are optimized, with multimodal prefix parameters $\Theta_f$ and Stage-I action-expert parameters $\psi$ frozen. Action-chunk generation:
$$a_t = G_{\psi,\theta}(z_t, \epsilon_t) \sim \pi_{\Theta_f,\psi,\theta}(\cdot\mid s_t), \qquad \epsilon_t\sim\mathcal{N}(0,I)$$
where $z_t = F_{\Theta_f}(o_t,\ell,q_t)$ is the task-conditioned multimodal prefix context and $G_{\psi,\theta}$ the flow-based action expert. The online RL objective:
$$\theta^\star = \arg\max_{\theta\in\mathbb{R}^p} J(\theta),\qquad J(\theta) = \mathbb{E}_{\tau\sim p(\tau\mid\pi_{\Theta_f,\psi,\theta},P,\rho_0)}\Big[\sum_{t=0}^{T-1}\bar{\gamma}^t R^{(H)}_t\Big]$$
ACoB Part 1: Progressive Value Calibration
ACoB instantiates the value model as an ensemble of $K$ task-specific critics, each using a value–advantage decomposition:
$$Q_{\phi_k}(\omega, \hat{a}) = V_{\phi_k}(\omega) + A_{\phi_k}(\omega, \hat{a})$$
where $\omega_t = (o_t, q_t) = \text{proj}_{o,q}(s_t)$ is the visual–proprioceptive component of the state, and $\hat{a} = \mathcal{C}(\mathcal{T}_t(a))$ the flattened critic representation.
Global return propagation. Let $a^{\text{exec}}_t$ be the action chunk ultimately executed — from autonomous inference or human correction — inducing transition $\xi_t = (s_t, a^{\text{exec}}_t, r_t, s_{t+1}, d_t)$. The bootstrap action and TD target are:
$$\hat{a}^{\theta_n}_{t+1} = \mathcal{C}\big(\tilde{G}_{\psi,\theta_n}(z_{t+1},\epsilon)\big),\qquad y_t = R^{(H)}_t + \bar{\gamma}(1-d_t)\min_k Q_{\bar{\phi}_k}(\omega_{t+1}, \hat{a}^{\theta_n}_{t+1})$$
with the ensemble TD objective:
$$\mathcal{L}_{\text{TD}}(\phi) = \mathbb{E}_{\mathcal{B}^{\text{RL}}_n}\Big[\frac{1}{K}\sum_k \big(Q_{\phi_k}(\omega_t, \hat{a}^{\text{exec}}_t) - \text{sg}(y_t)\big)^2\Big]$$
Recursive bootstrapping propagates long-horizon returns along executed trajectories.
Local preference ranking. The paper identifies a subtle but important gap: a correction can rescue a rollout to success, yet TD over the executed trajectory cannot correct the value of the overwritten proposal — which may remain overvalued and misguide action-expert optimization. ACoB adds state-matched local preference ranking. Before intervention the deployed VLA proposes $a^{\text{prop}}_t$; an effective intervention ($i_t=1$) supplies $a^{\text{hum}}_t$, giving:
$$a^{\text{cmd}}_t = \begin{cases} a^{\text{prop}}_t, & i_t = 0 \\ a^{\text{hum}}_t, & i_t = 1 \end{cases}, \qquad a^{\text{exec}}_t = \text{Exec}(s_t, a^{\text{cum}}_t)$$
with $c_t=1$ only when $i_t=1$ and $\lVert \hat{a}^{\text{exec}}_t - \hat{a}^{\text{prop}}_t \rVert_2 > \varepsilon_a$. Because the proposal has no observed successor, it is excluded from TD learning and used only in the ranking objective:
$$\mathcal{L}_{\text{rank}}(\phi) = \mathbb{E}_{\xi_t\sim\mathcal{B}^{\text{RL}}_n \mid c_t=1}\Big[\frac{1}{K}\sum_{k=1}^{K}\big[m_c - \Delta A^{\text{pair}}_{t,k}\big]_+^2\Big],\qquad \Delta A^{\text{pair}}_{t,k} = A_{\phi_k}(\omega_t, \hat{a}^{\text{exec}}_t) - A_{\phi_k}(\omega_t, \hat{a}^{\text{prop}}_t)$$
with margin $m_c\ge 0$ and $[x]_+=\max(x,0)$. The full critic objective:
$$\mathcal{L}_{\text{critic}} = \mathcal{L}_{\text{TD}} + \lambda_{\text{rank}} \mathcal{L}_{\text{rank}}$$
TD grounds the critic in long-horizon returns; local ranking directly calibrates the relative value of correction versus original proposal, resolving the local credit ambiguity that long-horizon TD cannot address. Under $Q=V+A$, $V$ is the shared state-only component while the ranking loss acts only on $A$ — isolating action-dependent differences from the state-only component and aligning the signal with action-expert optimization.
ACoB Part 2: Relative-Advantage Policy Improvement
Directly maximizing absolute $Q$ can amplify early value-estimation errors and drive the policy toward spuriously high-value actions. ACoB replaces absolute-value pursuit with relative-advantage improvement: paired advantage comparison focused on improvement between actions, reducing sensitivity to value scale and optimistic estimation error. But value-signal reliability must be established gradually, so this indirect signal alone gives slow, unstable early updates.
ACoB therefore couples relative-advantage improvement with flow-matching behavior cloning. Unlike prior methods that use BC mainly to constrain the policy distribution, ACoB uses critical-stage human corrections to drive BC early in training, rapidly absorbing high-quality corrections to improve the action expert and accelerate critic calibration through better experience. Long-horizon return propagation and local preference ranking then continually calibrate value estimates, providing increasingly reliable guidance as intervention recedes. This establishes asymmetric co-bootstrapping between rapid BC and progressive value calibration. Meanwhile the task-finetuned action expert is retained as a frozen reference: reference regularization keeps online refinement focused on execution precision rather than reshaping behavior, preserving competence and suppressing value-induced drift.
ACoB-Stream: Invariant-State Decoupling + On-Demand Streaming
To run ACoB on large VLAs, ACoB-Stream is a closed-loop experience–policy architecture organized around state lifecycles, with invariant-state decoupling and on-demand streaming as core design principles. The online experience pool comprises replay buffer $\mathcal{R}$, correction buffer $\mathcal{C}$, and context buffer $\mathcal{K}$: the actor generates $\tau_n$, storing all transitions in $\mathcal{R}$, effective corrections in $\mathcal{C}$, and multimodal prefix contexts in $\mathcal{K}$; the learner optimizes $\theta_n$ with ACoB to yield $\theta_{n+1}$ for redeployment:
$$\text{Actor}: \bar{\theta}_n \xrightarrow{\text{Rollout}} \tau_n \xrightarrow{\text{experience}} \mathcal{E}_n \ni \mathcal{B}_n,\qquad \text{Learner}: (\theta_n, \mathcal{B}_n) \xrightarrow{\text{ACoB}} \theta_{n+1} \xrightarrow{\text{policy}} \bar{\theta}_{n+1}$$
The paper is explicit about how this differs: unlike EXPO-FT and RL Token, which avoid propagating RL gradients through the VLA by optimizing residual action edits or a lightweight token-conditioned policy, ACoB-Stream lets improvements accumulate inside a unified policy; unlike fleet orchestration or cloud–edge scale-out, it reduces the per-update large-VLA computation itself. Measured gains reach 10.9× in throughput and computational efficiency.
Experiments
Hardware and Task Setup
Figure 3: Hardware platforms. Single-arm: one wrist camera and one external camera; dual-arm: two wrist cameras and one external camera.
Four platforms: (1) UR5e with a PGI-140-80 two-finger parallel gripper; (2) UR5e with a LinkerHand L20 dexterous hand; (3) two UR5e arms each with a PGI-140-80 gripper; (4) Franka Research 3 with a PGI-140-80 gripper. Nine high-precision chemistry-lab manipulation tasks span four categories: contact-rich, contact-light, contact-free, and bimanual coordination. To rigorously evaluate generalization, object poses and initial arm poses vary in every task. Training uses 4× NVIDIA A800 GPUs.
Main Results (9 chemistry manipulation tasks)
| Method | Mean training (min) | Mean success rate | Mean episode time (s) |
|---|---|---|---|
| HIL-SERL | — | 2.8% | 50.4 (×0.6) |
| ConRFT | — | 7.8% (+5%) | 45.3 (×0.7) |
| Robo-Dopamine | — | 10% (+7.2%) | 44.5 (×0.7) |
| π0 | — | 59.4% (+56.7%) | 32.0 (×1.0) |
| π0.5 | — | 67.8% (+65%) | 30.0 (×1.1) |
| VLA-Precision | 45.8 | 98.3% (+95.6%) | 27.6 (×1.2) |
VLA-Precision reaches 98.3% average success after 45.8 min of online training per task, with 27.6 s successful episodes at 1.2× and 1.8× the speed of VLA and RL baselines respectively. Per task: tube rack loading, 2 mL vial transfer, alcohol lamp extinguishing, pipette tip attachment, bulb dropper transfer, and pipette transfer/ejection all hit 100%; rubber stopper insertion 100% (vs HIL-SERL's 5%); cuvette transfer 90%; tube brushing 95% (vs π0 5%, π0.5 10%).
The baseline numbers are themselves informative: pure RL baselines nearly collapse on these precision tasks (HIL-SERL 2.8%, ConRFT 7.8%, Robo-Dopamine 10%), while pure VLA supervised fine-tuning is capped by demonstration quality (π0 59.4%, π0.5 67.8%). Together they confirm the paper's motivation: demonstrations cap the ceiling, and naive RL is unstable in real-world high-precision settings — the two must be coupled in the right way.
Shared Hyperparameters
| Hyperparameter | Value |
|---|---|
| $\lambda_{\text{rank}}$ | 50 |
| $(w_{\text{BC}}, w_{\text{rel}}, w_{\text{ref}})$ | (0.25, 0.50, 0.25) |
| Initial reset range | 3 cm in $x$, $y$, $z$ |
| Training GPUs | 4× NVIDIA A800 |
Significance and Limitations
demos → policy prior Θ_IL"] --> B["Stage II: Real-world online RL
only action-expert LoRA θ trainable"] B --> C["Actor rollout
autonomous or human-corrected"] C --> D["Pools: replay R + corrections C + context K"] D --> E["Critic: global return propagation L_TD"] C --> F["Critic: local preference ranking L_rank
correction > original proposal"] E --> G["Relative-advantage improvement
differences between actions, not absolute Q"] F --> G G --> H["Reference regularization
frozen Θ_IL suppresses drift"] H --> I["Deploy θ̄_{n+1} → rollout again"] I --> C
The most transferable idea is local preference ranking. It closes a gap that is easy to overlook: after a human intervention rescues a rollout, the overwritten proposed action was never executed and has no successor state, yet it may already be learned as high-value. Without explicitly pushing it down, that misestimate keeps misleading the policy. ACoB fixes this with a state-matched pairwise ranking (correction > proposal), and because of the $Q=V+A$ decomposition the ranking loss acts only on the advantage $A$, leaving the state value $V$ uncontaminated.
Three more engineering lessons. First, the "asymmetry" is across timescales — BC is fast, value calibration is slow; the former lifts performance early and feeds the latter better experience, then recedes as the latter takes over. This is more active than the common practice of merely using BC as a regularizer on RL. Second, freezing the Stage-I action expert as a reference confines online refinement to execution precision rather than behavior reshaping — a cheap and effective insurance against drift. Third, ACoB-Stream reduces the per-update large-VLA computation itself rather than scaling out with more robots and accelerators, making it orthogonal to — and composable with — cluster-scaling approaches.
On limitations, several boundaries are visible from the design and setup. The method depends on early human intervention: a human-in-the-loop teleoperation correction phase is required, and the quality and timing of those corrections directly shape subsequent bootstrapping. The task suite, while spanning four categories and four embodiments, is confined to benchtop-scale high-precision lab manipulation; transfer to longer-horizon mobile manipulation or more open settings is unverified. Hyperparameters such as $\lambda_{\text{rank}}=50$ and $(w_{\text{BC}},w_{\text{rel}},w_{\text{ref}})=(0.25,0.50,0.25)$ are shared across all tasks without a reported sensitivity analysis, and the influence of critic ensemble size $K$ and ranking margin $m_c$ is not ablated.
SOURCE LINKS



