PAPER DEEP DIVE
WB-WAM: Heterogeneous Body-Hand Pre-training for Humanoid Loco-Manipulation
WB-WAM from Tsinghua IIIS, Xiong'an Institute of AI and University of Melbourne (Hang Zhao group) injects explicit whole-body action supervision into generative video pre-training: a shared 72-D physical action space (body 29 + root 3 + two hands 20+20) unifies partial annotations from 1,880.2 hours of nine-source heterogeneous video and motion data via masked flow matching, followed by PICO egocentric mid-training (22 h / 73 tasks, GMR retargeting + constrained IK + MINT) and real-robot post-training with a forward-kinematics loss on Unitree G1 with Wuji hands. It wins all seven HumanoidArena tasks (81.9% mean), reaches 84.0% on five real tasks vs OpenWAM's 80.0% (ACT 20%, Fast-WAM 6%), and PICO mid-training lets 30 real demos hit 73.8% — beating 100-demo direct training at 65.0%, a 70% cut in robot data.
TL;DR
WB-WAM (Whole-Body World Action Model), from Tsinghua IIIS, the Xiong'an Institute of AI, and University of Melbourne (Hang Zhao's group; Chuan Qin and Shaoting Zhu equal contribution), injects explicit whole-body action supervision directly into generative video pre-training. A shared 72-D physical action space unifies body joints (29) + root motion (3) + two dexterous hands (20+20), letting 1,880.2 hours of partially annotated heterogeneous video and motion data train jointly; PICO egocentric mid-training and real-robot post-training with a forward-kinematics loss then land the model on Unitree G1 with Wuji hands. It wins all seven HumanoidArena tasks (81.9% mean), reaches 84.0% mean success on five real tasks vs OpenWAM's 80.0%, and PICO mid-training lets 30 real demonstrations hit 73.8% — beating 100-demo direct training at 65.0%, a 70% cut in robot data.
1. The Problem: Humanoid Loco-Manipulation Lacks Whole-Body Action Supervision
General-purpose humanoids must coordinate body motion and dexterous manipulation, yet large-scale robot pre-training data come mostly from single-arm, bimanual, and wheeled platforms — offering almost no direct supervision of coordinated humanoid body, root, and hand motion. Meanwhile heterogeneous motion sources are each incomplete: egocentric manipulation videos have fine hand annotations but no body motion; motion-capture collections have whole-body trajectories but no video. Combining them requires a prediction interface that accommodates partial annotation.
Existing systems ground actions differently — GR00T N1 uses flow matching, Ψ0 pre-trains on tokenized bimanual task-space actions, WholeBodyVLA learns discrete latent actions from visual transitions, OpenHLM adapts a non-humanoid VLA — but scalable video data either lack actions or use proxy representations. Directly incorporating temporally dense, physically interpretable body/root/articulated-hand references into large-scale world-action-model pre-training remained unexplored.
2. Method: a 72-D Shared Physical Action Space on a MoT Backbone
2.1 Action Space
Each whole-body action reference is a 72-dimensional physical vector:
$$a_t = \left[\, q^B_t,\; r_t,\; q^L_t,\; q^R_t \,\right] \in \mathbb{R}^{72}$$
with $q^B_t \in \mathbb{R}^{29}$ G1 body joint references, $r_t = (\phi_t, \theta_t, \omega^z_t) \in \mathbb{R}^3$ root roll, pitch, and yaw angular velocity, and $q^L_t, q^R_t \in \mathbb{R}^{20}$ articulated hand references. The space stays fixed across all three training stages — the key to letting heterogeneous sources "fill their own coordinates".
2.2 Architecture: Video Expert + Action Expert
Built on Fast-WAM's multimodal flow matching, the Mixture-of-Transformers backbone hosts a video expert and an action expert with modality-specific parameters; the action expert conditions on video features through MoT attention. Future RGB frames are encoded by a frozen video-model VAE; both experts model flow fields:
$$y_\tau = (1-\tau)\, y_0 + \tau\,\epsilon, \qquad u = \epsilon - y_0$$
The task is joint prediction $p_\theta(V_t, A_t \mid I_t, s_t, \ell)$ over future video and whole-body action trajectories given egocentric RGB, proprioception, and language. The total loss is:
$$\mathcal{L}(\theta; \mathcal{D}) = \lambda_{\text{vid}}\, \mathcal{L}_{\text{vid}} + \lambda_{\text{act}}\, \mathcal{L}_{\text{act}}$$
The crucial masked flow matching design: a binary mask $M$ marks annotated, non-padded action entries, and the target flow velocity applies only to valid channels:
$$u^a = M \odot (\epsilon^a - A_0)$$
So three supervision types — video+body+hand (D_VBH), video+hand (D_VH), text+body (D_TB) — train under one objective. During training, action tokens attend to either the full conditioning video or only the current frame with equal probability; at deployment the current-frame pathway denoises actions directly, without synthesizing future video.
2.3 How It Differs from Manipulation-Side WAMs
Three surgeries distinguish WB-WAM from fixed-base WAMs. First, the action target is whole-body joint references rather than end-effector poses — end-effector pose cannot distinguish "grabbing while standing" from "grabbing while crouching", and in humanoid tasks posture is part of the task (sitting on a sofa, boxing, ducking under a desk). Second, output splits into two paths: body/root references pass through a frozen SONIC encoder into a 64-D controller latent driving a 50 Hz whole-body tracking policy — outsourcing dynamic stability to a mature low-level stack — while the 40 hand dimensions bypass the encoder as direct joint commands for low-latency dexterity. Third, inference uses the current-frame pathway, decoupling action generation from video synthesis; replanning at ~0.7 Hz with 32-step chunks (1.6 s at 20 Hz), executing 20 steps per cycle, keeps motion continuous. Language and proprioception inject through UMT5 and a proprioceptive encoder via MoT attention — language modulates the joint denoising of both experts rather than directly emitting actions, which is what lets one policy serve the fruit-selection task.

3. Three-Stage Curriculum and WB-Datasets
| Stage | Data | Scale | Role |
|---|---|---|---|
| Stage I heterogeneous pre-training | Nine external sources (Xperience, EgoDex, HIW-500, MotionMillion, BONES-SEED, ...) | 1,880.2 h | Joint video-action learning; EgoDex 34.5% |
| Stage II PICO mid-training | PICO 4 Ultra + 5 trackers (GMR retargeting + constrained whole-body IK + MINT hand reconstruction) | 22 h / 13,396 episodes / 73 tasks | Task-aligned human motion supervision |
| Stage III real-robot post-training | SONIC teleoperation demos (MANUS gloves) | 1,011 episodes / 3.37 h / 8 tasks | Task adaptation + FK loss |
Stage III adapts a task-specific model $\theta^k_{\text{III}}$ per task, augmenting the joint video-action objective with a differentiable forward-kinematics loss (inspired by BeyondMimic's body tracking):
$$\mathcal{L}_{\text{III}} = \mathcal{L}(\theta; \mathcal{D}^k_R) + \lambda_{\text{FK}}\left(\mathcal{L}_{\text{pos}} + \beta_{\text{rot}}\, \mathcal{L}_{\text{rot}}\right)$$
Predicted action flows are denormalized into joint configurations and penalized on position/orientation errors of selected G1 links in the pelvis frame — binding "predicted references" to "kinematically reachable poses". The curation pipeline runs four checks: invalid poses, motion discontinuities (joint-command jumps), whole-body collisions, and corrupted/missing video frames; thresholded detector scores separate clear failures from ambiguous cases that get manual review.
3.1 The PICO Capture Pipeline
Mid-training data capture uses a PICO 4 Ultra headset plus five trackers (hands, feet, waist), recording egocentric RGB and 20 Hz SMPL body motion. The conversion chain: GMR produces an initial G1 motion sequence; constrained whole-body inverse kinematics refines it — tracking calibrated human task-space positions and palm orientations while respecting joint limits, velocity bounds, support constraints, and self-collision avoidance; MINT reconstructs human-hand keypoints from the egocentric video and retargets them to Wuji joint configurations; finally body/root references are temporally aligned with the recorded video. The authors are candid that retargeting provides only kinematic supervision — dynamic stability and contact feasibility are not established by retargeting alone, which is exactly why Stage III still needs real demonstrations and the FK loss.

4. Results
4.1 Simulation: Winning All Seven HumanoidArena Tasks
| Method | Football | DoubleDesk | P&PBox | OpenDoor | SitSofa | Boxing | VisNavi | Mean |
|---|---|---|---|---|---|---|---|---|
| Best baseline (per task) | 45.0 | 43.3 | 75.0 | 85.0 | 78.3 | 76.7 | 38.3 | — |
| π0.5 / DP / ACT / FM | 51.2 | 47.6 | 45.0 | 60.0 | 51.2 | 60.0 | — | — |
| WB-WAM | 70.0 | 65.0 | 86.7 | 98.3 | 95.0 | 81.7 | 76.7 | 81.9 |
Gains span locomotion (Football +25.0, VisNavi +38.4), posture adjustment (SitSofa +16.7), and object interaction (P&PBox +11.7) — precisely where whole-body pre-training differs from manipulation-only pre-training.
4.2 Real World: Six Baselines on G1
Five tasks (wipe the table, close the curtain, make the bed, move the pillow, tidy the cloth; the first three involve locomotion), 20 trials each, ~100 demos per task, against ACT, π0.5, GR00T N1.6, Fast-WAM, DiT4DiT, and OpenWAM. WB-WAM reaches 84.0% mean SR (86.67% mean TP) vs OpenWAM's 80.0%; ACT manages only 20% (0/20 across all locomotion tasks) and Fast-WAM 6%. On locomotion tasks WB-WAM leads 76.7% vs 71.7%. Baseline failures typically involve inaccurate approach, failed grasping, or poor body-manipulation coordination. Per-task, move-the-pillow saturates at 100% for both WB-WAM and OpenWAM; wipe-the-table is the one loss (75% vs 90%) — a task demanding sustained forward lean and stable downward force, suggesting chunk-boundary consistency for long-horizon contact still has headroom. Milestone decomposition (Approach/Contact/Success) shows WB-WAM's edge concentrated in the C→S segment — coordinated execution after contact.
4.3 Language-Conditioned Manipulation and Generalization
A single policy trained on ~300 demonstrations spanning orange, lemon, and apple achieves 75–85% contact and success per fruit (15/20, 17/20, 16/20) without per-target training — language directly modulates body and hand trajectory generation. In the visuomotor generalization study, extra objects alter the cart's appearance while the objective stays fixed: the video expert rolls out future views under the modified condition and WB-WAM executes real-robot cart pushing without any adaptation — the video expert's prediction capability buffers OOD conditions for the action expert, a structural advantage of joint video-action modeling over pure action policies.
4.4 PICO Mid-Training: 70% Fewer Robot Demos, Higher Success
| Training route | Robot demos | Mean SR | Mean TP |
|---|---|---|---|
| Direct post-training | 30/task | 48.8% | 51.0% |
| Direct post-training | 100/task | 65.0% | 70.4% |
| PICO mid-training + post-training | 30/task | 73.8% | 78.1% |
With ~150 task-aligned PICO demos per task as mid-training, 30 robot demos (a 70% cut) reach 73.8%, beating 100-demo direct training at 65.0%. Mid-training also raises the ceiling: on overlapping tasks it scores 90% (wipe) and 80% (curtain) vs the no-PICO versions' lower marks — cheap VR-recorded human demonstrations carry genuine task-aligned priors. The cleanest ablation in the paper is Fast-WAM 6% vs WB-WAM 84%: same MoT backbone, same Wan2.2 initialization — the 78-point gap is almost entirely attributable to "body and hands in pre-training". Meanwhile the three non-whole-body baselines show a sharp w/-locomotion vs w/o-locomotion gap that WB-WAM largely closes. Reproduction cost structure: the 1,880 h of external data is nearly free (public datasets + retargeting); the 22 h of PICO capture needs only a headset and five trackers; the 1,011 real teleoperation episodes are the only expensive part — and the PICO route cuts even that by 70%.
video+body+hand / video+hand / text+body"] --> B["Stage I: heterogeneous pre-training
MoT video + action experts
masked flow matching"] B --> C["Stage II: PICO mid-training
22 h / 13,396 eps / 73 tasks
GMR retarget + constrained IK + MINT"] C --> D["Stage III: real-robot post-training
SONIC teleop demos
+ forward-kinematics FK loss"] D --> E["Deployment: Unitree G1 + Wuji hands
SONIC 64-D latent for body
direct joint commands for hands"] E --> F["HumanoidArena 81.9%
all seven tasks won"] E --> G["Real five tasks 84.0%
vs OpenWAM 80.0%"] E --> H["PICO transfer: 30 demos 73.8%
beats 100-demo direct 65.0%"]
5. Limitations and Outlook
The authors note the evaluation is limited to one robot platform and task-specific policies (one $\theta^k_{\text{III}}$ per task); extending transfer to unseen tasks and improving execution accuracy during object interaction are future work. Retargeting provides kinematic supervision only — dynamic stability and contact feasibility still require real-robot post-training. Strategically, WB-WAM answers "how does a WAM go whole-body" (supervision coverage) while Rolling-WAM answers "how does a WAM go fast" (rolling denoising); the PICO result delivers a plain but powerful engineering conclusion: task-aligned human demonstrations are the highest-value substitute for real-robot data — 22 hours of VR-level recording offsets 70% of robot collection cost. Project page: wb-wam.github.io.
6. Positioning: Three Lines Converging in 72 Dimensions
WB-WAM sits at the intersection of three 2026 threads: World Action Models supply the joint video-action skeleton (Fast-WAM's MoT flow matching, Rolling-WAM's rolling denoising); humanoid foundation models supply the heterogeneous human-robot pre-training paradigm (GR00T N1, Ψ0, WholeBodyVLA); human motion transfer supplies the path around the robot-data bottleneck. Its unique contribution is fusing all three inside one 72-D physical action space with masked flow matching — not normalizing the data into one format, but letting the action space host each source's native format. Compared with GE-Act 2.0 (sharing generative video priors) and UCAG-P (sharing camera-observable geometry), WB-WAM shares physical joint-space coordinate channels — body, root, and hand each keep their slot, and whoever has labels fills it. Where UCAG-P answers "where is the hand and how does it grasp", WB-WAM answers "how does the body get there and stand". A natural next step is unifying camera-centric anchor geometry with 72-D whole-body references — at which point human videos, robot trajectories, and mocap would participate in humanoid pre-training without distinction.
SOURCE LINKS

