PAPER DEEP DIVE
World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal
HKUST(GZ), CUHK and Knowin AI present World Action Agent (WAA), a multi-agent harness that changes what the VLM sees and how its decisions take effect — rather than the VLM itself — so a general-purpose VLM can pilot a robot with basic tools (point, drag, preview, execute), making every decision inside one visual action workspace. Three properties: Contact views, whose camera parameters are solved from interaction-region visibility under occlusion, framing compactness, view redundancy and stability, with the feasible set constraining the two views to orthogonal horizontal projections so any alignment error can be read along two independent directions — requiring only a base-frame point cloud, so the perception source (simulation, fused RGB-D, or VGGT reconstruction) is irrelevant; action rehearsal, where each action is an editable proposal planned by cuRobo, overlaid as a translucent robot in every view with a feasibility report, refined by an Imagination Agent in a separate context while the physical scene stays unchanged, and only proposals with executable plans become motion; and in-view correction, where the agent drags from a reference point to the desired location in the Contact view where the error is observed — the reference can be a visible point on the held object, so no conversion into an absolute gripper pose is needed. Skills evolve from expert videos and human Canvas-GUI teaching through Learner, Editor and Reviewer roles under evidence-citing and independent-review constraints. On LIBERO-Pro with Gemini 3.7 Flash, using skills evolved only from LIBERO-90 and frozen before evaluation, WAA reaches a state-of-the-art 75.6% average success, above ASPIRE (72.0%), end-to-end VLAs and the same-backbone Show-Harness (6.7%), scoring 80.0 / 73.3 on the two Spatial splits; an episode costs just 31 model calls, 150 s and $0.1996 versus 120 calls, 874 s and $0.5021 for Show-Harness. The frozen skills transfer to robosuite at 100.0%, and LoRA fine-tuning Qwen3.5-9B on 112 trajectories (1,774 decision steps) as the main agent only raises out-of-domain success from 1.7% to 43.3%.
Source: arXiv:2609.29964 (cs.RO) — HKUST(GZ), CUHK, and Knowin AI. Submitted Sep 24, 2026; 29 pages. Marked by the authors as work in progress.
In One Sentence
General-purpose VLMs have broad knowledge and spatial reasoning, but existing systems use them at arm's length — predicting constraints, writing programs, or showing them a picture of the scene rather than a world in which to act. World Action Agent (WAA) argues: don't change the VLM, change what it sees and how its decisions take effect. It builds a multi-agent harness in which a VLM pilots a robot using only basic tools (point, drag, preview, execute), making every decision inside one visual action workspace. That workspace has three properties: automatically selected Contact views center the camera on the current interaction; action rehearsal turns each action into an editable proposal that is previewed and revised against planning feedback before execution; and in-view correction lets the agent remove residual error in the very view where it observes it. On LIBERO-Pro, with skills evolved only from LIBERO-90, WAA reaches 75.6% average success (state of the art), beating end-to-end VLAs, code-as-policy agents, and a visual-harness baseline on the same backbone; the same frozen skills transfer to robosuite without further learning; and fine-tuning Qwen3.5-9B on harness traces lifts its out-of-domain success from 1.7% to 43.3%.
1. Framing: How Should a VLM Actually Get Its Hands on a Robot?
1.1 What's Missing in Each Existing Route
- End-to-end VLAs (RT-2, OpenVLA, $\pi_{0.5}$): fine-tune the VLM into an action predictor. The problem — fine-tuning VLMs for action prediction may weaken their general understanding and reasoning (Hancock et al., 2026).
- World action models (WAMs): couple environment dynamics with action learning. But grounding learned visual dynamics in executable control still needs robot data for action alignment, and cross-embodiment transfer stays sensitive to morphology and action-space differences.
- Constraint / program routes (HAMSTER, ReKep, CaP, CaP-X): have the VLM predict intermediate paths, relational constraints, or write programs. The VLM participates only indirectly — it supplies constraints or code rather than operating robot primitives through its own spatial understanding.
1.2 The Three Things Visual Harnesses Still Lack
Show-Harness and VIA take a step forward: the VLM selects and revises actions directly through a visual robot interface, turning manipulation into interactive visual reasoning. But the authors make a sharp point — such interfaces display the scene to the VLM and let it choose actions; they do not give it a world in which to act. Three things are missing:
- Observation is not centered on the interaction. Success hinges on local relations among gripper, object and target, yet a fixed global camera often cannot reveal them because of distance, occlusion, or degenerate viewing angles.
- Actions cannot be tried before they are taken. Each takes effect as soon as it is issued, before the VLM can see its consequences.
- Perception and action live in different spaces. An offset observed in an image must be rewritten as coordinates before it can be corrected.
So the question is reframed: how can a general-purpose VLM use basic action primitives to make and revise decisions throughout execution, serving as a robot pilot? WAA's answer is to change what the VLM sees and how its decisions take effect — not the VLM itself.
Figure 1: WAA overview — visual action rehearsal plus reusable skills let VLMs perform diverse manipulation tasks and adapt through demonstrations and human guidance.
2. Method: The Visual Action Workspace
2.1 Formulation
Given task instruction $g$, the agent completes the task through a sequence of tool calls. At step $t$ the harness presents a multi-view Canvas $C_t$ and a control context $m_t$ recording valid spatial references, the pending action proposal $\hat{a}_t$, and recent execution feedback. The policy selects a tool call:
$$u_t \sim \pi_\theta(\cdot \mid g, C_t, m_t)$$
where $u_t$ specifies a tool and its arguments. Tools fall into three classes by effect:
- query tools — acquire spatial information or skill guidance;
- proposal tools — construct and revise $\hat{a}_t$ and return a visual preview with planning feedback;
- execution tools — actually move the robot.
The key design: query and proposal tools leave the physical world unchanged, so the agent can inspect and revise an action repeatedly before committing. The VLM decides what to manipulate, how to revise a proposal, and when to execute; the harness turns those decisions into geometrically accurate, kinematically feasible motion.
Global + Contact views"] --> M["Main Agent
decides: query / proposal / execute"] M -->|needs spatial refinement| I["Imagination Agent
iterates on the proposal in a separate context
physical scene unchanged"] M -->|needs procedural knowledge| S["Skill Agent
compares reference images with current Canvas
reports applicability and differences"] I -->|returns revised proposal| P["Editable proposal + visual preview
+ planning feasibility feedback"] S -->|skill guidance| P P -->|"only proposals with executable plans"| E["Execute: move / release / regrasp"] E -->|new observation + feedback| O E -.->|residual offset| C["In-view correction
drag in the view where error is seen"] C --> O
Figure 2: WAA harness overview — (A) the main agent's visual action workspace, (B) skill learning and evolution, (C) training a smaller VLM on interaction traces.
2.2 Design One: Interaction-Centered Canvas and Contact Views
Outcomes often hinge on millimeter- to centimeter-scale relations among gripper, object and target that a fixed global camera frequently fails to reveal. WAA therefore selects viewpoints actively for the current interaction. Given scene point cloud $P_t$, robot geometry $R_t$, and highlighted interaction region $H_t$, the harness solves for Contact view camera parameters:
$$c_t = \arg\max_{c \in \Omega} \left[ S_{\text{vis}}(c; H_t, R_t \mid P_t) + \lambda_1 S_{\text{frame}}(c; H_t) - \lambda_2 E_{\text{red}}(c) - \lambda_3 E_{\text{stab}}(c, c_{t-1}) \right]$$
The four terms:
- $S_{\text{vis}}$ — visibility of the interaction region and gripper under scene occlusion;
- $S_{\text{frame}}$ — encourages compact framing;
- $E_{\text{red}}$ — discourages views conveying the same direction;
- $E_{\text{stab}}$ — suppresses viewpoint jumps so spatial relations stay comparable across steps.
And the feasible set $\Omega$ carries a deliberate constraint: the horizontal projections of the two Contact views must be orthogonal. This maps directly onto the third gap in §1.2 — orthogonality means any alignment error can be read along two independent directions, so no offset can hide in a viewing blind spot.
The objective only requires a scene point cloud in the robot base frame, so WAA is agnostic to how $P_t$ is obtained: rendered in simulation, fused from calibrated RGB-D cameras, or reconstructed from RGB with a feed-forward model such as VGGT. All three are tested. Every view retains its projection calibration, so image-space annotations map directly to 3D — the precondition that makes in-view correction possible.
2.3 Design Two: Action Rehearsal
For a VLM the hard part is often not deciding where to go but judging whether a specific pose is appropriate: does the approach collide with nearby objects, can the robot reach it, does the resulting grasp support the next step? WAA therefore represents an action as an editable proposal rather than a one-shot output:
- The agent localizes targets on the Canvas and specifies a target pose;
- the harness solves inverse kinematics and plans motion with cuRobo;
- it overlays the target configuration as a translucent robot in every view;
- and reports feasibility.
The agent can thus examine approach direction, prospective contact and surrounding clearance before execution, and revise accordingly. For finer adjustment the main agent delegates spatial intent to an Imagination Agent running in a separate context: it iteratively edits and replans while the physical scene remains unchanged, returning only the revised proposal. The hard rule: only a proposal with an executable plan can become physical motion.
The paper's example is concrete: an initial pose yields "Motion Plan Infeasible"; the Imagination Agent adjusts via rotate(axis, deg) until planning succeeds; the main agent then shifts it toward the handle.
Figure 3: Action rehearsal and in-view correction — (a) an infeasible pose becomes feasible; (b) Contact A appears aligned while Contact B reveals the offset.
2.4 Design Three: Closed-Loop Execution and In-View Correction
Target-level planned motions suit large movements like approaching and transporting, but depth noise, calibration error, and contact disturbances leave residual errors that often decide whether placement succeeds. WAA lets the agent express corrections directly in the view where the error is observed: it drags from a reference point to the desired location in a Contact view, and the harness uses that view's calibration to convert the image-space intent into a bounded end-effector displacement.
One easily overlooked but crucial detail: the reference point can be the gripper or a visible point on the held object. The agent can state where the object should go without first converting that relation into an absolute gripper pose — a direct answer to the third gap in §1.2. Two orthogonal Contact views support corrections along complementary directions. After each correction a fresh observation lets the agent decide whether to adjust further, release, or proceed.
Figure 3b shows why this matters: Contact A appears aligned, but Contact B reveals an offset; the main agent drags the bowl toward its target in Contact B, the harness converts the drag into end-effector motion, and the bowls end up stacked.
3. Learning Through the Harness
The workspace lets a VLM observe and act on spatial relations, but does not tell it how to use those capabilities to complete a task — that needs embodied procedural knowledge VLM pretraining rarely captures: which contact to establish first, which relation to verify before release, how to recover from a failed grasp. WAA acquires it along two complementary paths.
3.1 The Multimodal Skill Library
Textual procedures alone struggle to convey judgments like "is this alignment sufficient," while visual references supply the evidence needed for state recognition and verification. Each WAA skill therefore combines applicability conditions, a procedure, key states with reference images, and observable outcome checks. The procedure specifies the relations the gripper, object and target should satisfy, the required contact and clearance, and how to recover from failure — rather than fixed coordinates or displacements. Objects, target positions and motion magnitudes are always resolved in the current scene; numerical values appear only as tentative ranges to adjust from observed feedback.
That is the essential difference from libraries like ASPIRE's, which register validated per-object parameters (grasp height offsets, yaw angles). WAA's relational guidance transfers across object instances, layouts and viewpoints. During execution the main agent selects a skill by name and description and delegates interpretation to a Skill Agent in a separate context: it selects relevant procedural states and reference images, compares them with the current Canvas, and reports applicability, scene differences, and relations it cannot yet confirm.
A nice engineering detail: the main agent receives the full textual procedure while reference images stay in the Skill Agent's context — preventing historical images from being mistaken for the current scene and keeping the main context compact. Scene-specific advice expires when its associated observation or control state changes, so the agent never acts on outdated guidance.
3.2 Evolution from Demonstrations and Teaching: Learner / Editor / Reviewer
The library starts from text-only seed skills that Codex (GPT-5.5) writes from the harness API, then is enriched with one expert trajectory per LIBERO-90 source task (scenes differing from evaluation). Learning proceeds through three GPT-5.5-driven roles:
| Role | Job | Key constraint |
|---|---|---|
| Learner | Segments each video at observed state changes, interprets contact, object motion and support changes as harness actions, extracts knowledge candidates | Does not read existing skill text, so new evidence is not steered by prior conclusions; candidates must cite source frames |
| Editor | Compares candidates with existing procedures and reference images; adds, revises, retires or defers using source frames as visual evidence | Every edit must carry evidence |
| Reviewer | Independently checks revised skills against their source evidence and returns concrete issues to the Editor | Only supported changes that preserve existing applicability conditions enter the library |
Demonstrations show successful behavior and rarely reveal recovery from mistakes. WAA therefore provides a Canvas-GUI: humans see the same multi-view Canvas and operate the robot with the same tools — pointing in a view, dragging in a Contact view, previewing and executing proposals. Humans pilot the robot much as a GUI agent would, and the harness records these actions in exactly the same format as agent interactions. The Learner contrasts the agent's attempt before failure with the human correction from the same state, identifies which relation (contact location, alignment direction, release timing) accounts for the outcome difference, and distills recovery guidance through the same Editor–Reviewer loop.
3.3 Learning to Pilot: Distilling Traces into a Small Model
Large proprietary VLMs pilot the harness well but at high cost and latency. WAA distills interaction traces into a smaller VLM operating the same harness. Training examples come from successful episodes (only clean executions without redundant operations), plus separately reviewed recovery segments such as regrasping after an empty grasp. Each example pairs the pre-action Canvas, task instruction, control context and available skill guidance with the executed tool call, without the teacher's reasoning or any privileged simulator state. Training is maximum likelihood on the policy in Eq. (1):
$$\max_\theta \sum_{(g, C_t, m_t, u_t) \in \mathcal{D}} \log \pi_\theta(u_t \mid g, C_t, m_t)$$
Crucial boundary: fine-tuning updates only the main policy while the workspace and sub-agents stay fixed — the model learns to use the interface, not to replace it.
4. Experiments
4.1 Setup
Benchmarks: LIBERO-Pro (Object / Goal / Spatial splits, each under initial-position swaps (Pos.) and task perturbations (Task), ten tasks per split; 60 episodes per split, six per task) and robosuite (cube lifting / stacking / restacking). Backbone Gemini 3.7 Flash in three settings: no skills (zero-shot), text-only seed skills, and the evolved library. The skill library is learned only from LIBERO-90 and frozen before evaluation — no LIBERO-Pro or robosuite rollout, failure log or label updates skills, prompts or parameters. Each episode is limited to 50 main-agent turns, 50 physical operations and one hour, with at most six turns per Imagination Agent call.
4.2 Main Results: LIBERO-Pro
| Method | Class | Obj. Pos. | Obj. Task | Goal Pos. | Goal Task | Spa. Pos. | Spa. Task | Avg. |
|---|---|---|---|---|---|---|---|---|
| OpenVLA | End-to-end VLA | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| $\pi_{0.5}$ | End-to-end VLA | 17.0 | 1.0 | 38.0 | 0.0 | 20.0 | 1.0 | 12.8 |
| CaP-Agent0 (CaP-X) | Code-as-policy | 22.0 | 18.0 | 26.0 | 17.0 | 12.0 | 14.0 | 18.2 |
| CaP-Agent0 (Playful) | Code-as-policy | 27.0 | 31.0 | 29.0 | 16.0 | 13.0 | 23.0 | 23.2 |
| RATs (Playful) | Code-as-policy | 61.0 | 63.0 | 43.0 | 36.0 | 29.0 | 31.0 | 43.8 |
| ASPIRE | Code-as-policy | 98.0 | 95.0 | 81.0 | 45.0 | 51.0 | 60.0 | 72.0 |
| Show-Harness | Visual harness | 13.3 | 16.7 | 0.0 | 10.0 | 0.0 | 0.0 | 6.7 |
| WAA (zero-shot) | Ours | 53.3 | 46.7 | 30.0 | 23.3 | 13.3 | 6.7 | 28.9 |
| WAA + seed skills | Ours | 83.3 | 70.0 | 26.7 | 40.0 | 16.7 | 23.3 | 43.3 |
| WAA (evolved skills) | Ours | 95.0 | 93.3 | 63.3 | 48.3 | 80.0 | 73.3 | 75.6 |
| WAA (evolved + RGB-D) | Ours | 88.3 | 91.7 | 56.7 | 45.0 | 76.7 | 68.3 | 71.1 |
| WAA (evolved + VGGT) | Ours | 90.0 | 88.3 | 53.3 | 45.0 | 65.0 | 71.7 | 68.9 |
Four readings:
- 75.6% is state of the art, above ASPIRE's 72.0%, with skills learned only from LIBERO-90 and frozen before evaluation.
- The largest margins are on the Spatial splits (80.0 / 73.3), where success depends on centimeter-scale placement relations; ASPIRE remains strongest on both Object splits and Goal Pos.
- Under a fair same-backbone, same-no-skills comparison, Show-Harness scores only 6.7% (failing every Spatial episode) versus 28.9% for zero-shot WAA. Both let the VLM select and revise actions; the difference is what the agent can observe and test before acting.
- End-to-end VLAs nearly collapse under task perturbations ($\pi_{0.5}$ falls from 17.0 to 1.0) as unfamiliar instructions and object combinations fall outside their training distribution, whereas WAA preserves the VLM's general understanding and grounds each decision in the current Canvas.
4.3 Cost: Cheaper and Much Faster
| Method | API cost ($) ↓ | Model calls ↓ | Time (s) ↓ | Input tokens (K) ↓ | Output tokens (K) ↓ |
|---|---|---|---|---|---|
| Show-Harness | 0.5021 | 119.80 | 874.12 | 347.10 | 62.90 |
| WAA | 0.1996 | 31.05 | 150.47 | 193.62 | 5.45 |
31 model calls and 150 s per episode against 120 calls and 874 s, with one twelfth the output tokens. The authors credit expressing spatial intent as in-view drags rather than long coordinate reasoning, and the Imagination Agent absorbing much trial-and-error in a separate context.
Figure 4: Qualitative WAA examples — (A) drag-guided placement, (B) rehearse and revise, (C) revised contact enables a successful lift.
4.4 Skill Learning, Transfer and Generality
- Evolved multimodal skills contribute the largest gain: 28.9% (no skills) → 43.3% (seed skills) → 75.6% (evolved); Spatial rises from 13.3 / 6.7 to 80.0 / 73.3. Text-only seed skills help inconsistently and even slightly lower Goal Pos., while evolved skills improve on zero-shot in all six splits.
- One demonstration suffices for a new skill. Starting from seed skills with no stove-specific procedure, a single LIBERO-90 demonstration yields a skill with which WAA succeeds in all ten executions of the LIBERO-Pro stove task. Notably, the Reviewer rejects the first draft for treating knob motion as success; the accepted skill requires an explicit activation signal — evidence that independent review blocks overly broad outcome checks.
- LIBERO-learned skills transfer to robosuite, whose scenes, objects and camera placements all differ: without skills WAA already succeeds in every lifting and stacking trial and reaches 60.0% on restacking; the frozen LIBERO skills raise restacking to 100.0% (100.0% average) against 63.3% for CaP-Agent0 with RATs skills.
- Insensitive to point-cloud source. Swapping the simulator cloud for fused RGB-D or VGGT reconstruction lowers the average only to 71.1% and 68.9% (−4.5 and −6.7 points), and both still surpass every baseline on the two Spatial splits. Because Contact-view selection only needs a base-frame point cloud, the perception source can change without modifying agents or skills.
- A smaller VLM learns to pilot the same harness. Fine-tuning Qwen3.5-9B with LoRA (rank 8, 10 epochs, lr $5\times10^{-5}$) on 112 harness trajectories (1,774 decision steps), replacing only the main agent while sub-agents and grounding stay on Gemini 3.7 Flash: in-domain 0.0% → 55.0%, out-of-domain 1.7% → 43.3%.
5. Our Take
The real contribution here is not the metric but treating interface design as the object of study. WAA changes no model and trains no policy; it redesigns the layer between VLM and robot — and that layer is exactly what decides whether the VLM's spatial understanding cashes out into executable motion. The most telling thing about 75.6% is not that it beats ASPIRE, but that with the same backbone and likewise without skills, Show-Harness gets 6.7% while zero-shot WAA gets 28.9%: the difference is those three designs.
Of the three, the one most worth remembering is in-view correction, especially the fact that the reference point can be a visible point on the held object. It frees the agent from having to convert an observed relation into an absolute pose first — correct the error in the view where you see it, putting perception and action in the same space. Combined with the constraint that the two Contact views must have orthogonal horizontal projections, the interface guarantees at a geometric level that errors cannot hide in a viewing blind spot. That is clean geometric thinking.
Other points worth singling out:
- The Learner does not read existing skill text — a counterintuitive constraint that stops new evidence being steered by prior conclusions; with the Reviewer's independent check, it is a concrete engineering answer to the familiar problem of skill libraries degrading as they learn.
- Reference images stay in the Skill Agent's context, not the main context, so historical images are never mistaken for the current scene — a very practical pitfall in multimodal agent systems.
- The Canvas-GUI lets humans operate the robot with exactly the same tools, so human corrections and agent traces share a format and serve as supervision naturally, with no extra data pipeline.
- Only the main policy is fine-tuned; the interface never changes, so the model learns to use the interface rather than replace it — meaning interface improvements keep benefiting every pilot.
Limitations are stated plainly: performance remains bounded by the backbone — Gemini 3.7 Flash still integrates multi-view information imperfectly on some tasks, making WAA less stable there. The authors note this lies in the backbone's perception rather than the harness interface, so stronger multi-view VLMs benefit WAA directly. On cost: WAA closes the loop more frequently than code-as-policy methods, grounding every decision in a new observation; the Flash-class backbone is cheaper than frontier proprietary models, yet an episode still requires tens of model calls, making smaller trained pilots the main route to lower cost and latency.
One caveat for readers: the paper is marked work in progress, and among the LIBERO-Pro comparisons only Show-Harness was rerun by the authors (30 episodes per split) — other baseline numbers are taken from the original papers, which is worth keeping in mind when comparing across papers.
6. Resources
- Paper: arXiv:2609.29964
- Affiliations: HKUST(GZ), CUHK, Knowin AI
- Status: Work in progress
SOURCE LINKS