PAPER DEEP DIVE
EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data
Egocentric human data offers scalable supervision for robot manipulation. However, behavior cloning entangles transferable content like objects, scenes, and task semantics, with non-transferable factors like human morphology, head motion, and behavioral style. We study whether World Action Models (WAMs) provide a better training signal by requiring policies to predict not only actions, but also how the scene evolves. The central question is what world representation best enables human-to-robot transfer. We hypothesize that an effective world target should abstract appearance, capture agent-invariant physical effects, and separate camera motion from environment change. We introduce EgoWAM, a controlled human-robot co-training framework that fixes the policy backbone, action head, and data mixture while varying only the world prediction target, comparing Pixel, DINO, and 3D motion flow. Across three real-world bimanual tasks, WAM co-training scales more effectively with in-the-wild egocentric human data than behavior cloning. Pixel-based prediction transfers weakly, while DINO and 3D flow yield substantial gains: DINO improves out-of-distribution object and scene generalization by up to 4x, and 3D flow improves in-domain performance by 20-30%. More details: https://gatech-rl2.github.io/egowam.github.io
Paper: EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data
Authors: Baoyu Li*, Xinchen Yin*, Mengying Lin, Yixin Zhang, Danfei Xu (Robot Learning and Reasoning Lab, Georgia Institute of Technology; * equal contribution, correspondence bli678@gatech.edu)
Links: arXiv:2607.08436 (cs.RO, 8 Jul 2026, v1, CC BY 4.0, CoRL 2026) · project page gatech-rl2.github.io/egowam.github.io
Code: released, GaTech-RL2/EgoWAM (MIT license, repository description tagged CoRL 2026, 17 stars at the time of writing, last push 2026-09-07).

Figure 1: The opening figure. Lab robot data and in-the-wild egocentric human data are separated by an embodiment gap made of camera motion, morphology, and behavioral style. (1) BC co-training has only the action head as a channel, so the gap leaks into the policy as motions the robot cannot execute; WAM co-training opens a second channel that predicts scene evolution, transferring human data through dynamics where actions cannot go. (2) The question EgoWAM studies: which world representation actually enables transfer, judged against appearance abstraction, cross-embodiment consistency, and ego-motion factoring.
One-Line Summary
A controlled human-robot co-training study that freezes the backbone, the action head, and the data mixture and varies only what the world-model head is asked to predict: pixel VAE latents barely transfer, DINO features lift out-of-distribution object and scene generalization by up to 4x, camera-stabilized 3D flow lifts in-domain scores by 20-30%, and the world-model head is discarded at inference, so deployment costs exactly what a same-size BC policy costs (30 Hz on one RTX 4090).
Background and Motivation
Data is the binding constraint in robot manipulation. A single bimanual teleoperated demonstration costs a human operator tens of seconds of undivided attention, while the space of objects, scenes, and task combinations diverges. Recording a person doing household work through egocentric glasses such as Project Aria is an order of magnitude cheaper to collect and carries object and scene diversity that no lab bench can reproduce. Scaling robot manipulation on human video is therefore one of the most actively pursued routes in embodied AI.
The most direct way to use that video is behavior-cloning co-training: retarget human demonstrations into a robot-executable action space, mix them with robot demonstrations in the same batch, and train one imitation policy on both. Recent work shows this paradigm scales, but with a condition attached. The human data has to be carefully aligned to the robot in viewpoint, motion speed, and behavioral style. Without that bridge, action-level co-training injects human-like motions the robot cannot execute and degrades downstream performance.
The authors name this failure mode the bitter lesson of action-level co-training. The shared action decoder is the sole channel through which human data reaches the policy, so it is forced to entangle transferable content (objects, scenes, task semantics) with non-transferable execution (morphology, behavioral style), and the embodiment gap living in the latter blocks the transfer of the former. This is not something hyperparameter tuning repairs; it is a structural property of the interface.
World Action Models open a second supervision channel. An auxiliary world-model head reads the same shared latent and predicts a future state, grounding action prediction in task-relevant dynamics. The decisive property is that this channel operates on observations rather than actions: how a scene is about to change depends far less on whether a human hand or a parallel-jaw gripper caused the change. The paper's central hypothesis follows directly. Task-relevant dynamics transfer across embodiments more readily than actions do, so human data can shape the shared backbone through the world-model channel even when its action labels cannot be trusted.
That promise rests on a question prior WAM work never studied systematically: predict what? Most existing WAMs predict pixels through a pretrained video VAE whose latent is optimized for photometric reconstruction and entangles motion with appearance. The paper identifies that entanglement as the dominant failure mode of pixel-level WAM co-training, and proposes three desiderata for the world target. (D1) Appearance abstraction: targets that reward photometric reconstruction force the trunk to encode embodiment-specific appearance, crowding out the structure that governs task outcomes. (D2) Cross-embodiment consistency: targets should represent the effect rather than the agent, so that a human hand and a robot gripper producing similar physical change induce similar supervision. (D3) Ego-motion factoring: image-coordinate targets conflate head rotation with scene change, giving the same physical event different supervision under a moving human camera versus a static robot camera.
Two alternatives are then studied along that axis. DINO features abstract appearance through a semantic prior and align predictions across embodiments, but stay spatially indexed on the image grid, so (D3) is only partially mitigated. 3D motion flow satisfies all three desiderata by construction through camera-frame-aligned geometric grounding. EgoWAM is the framework that puts all three targets inside one training loop so the comparison is apples-to-apples.
Preliminaries
The HPT backbone. EgoWAM builds on the Heterogeneous Pretrained Transformer, where embodiment-specific stems tokenize each input into a shared latent space and a single transformer trunk operates on the unified token stream. That inductive bias extends naturally to world action models: observation tokens, action tokens, and future-prediction tokens coexist in one stream and attend within the same trunk, which turns future dynamics and action prediction into two read-outs of one shared latent rather than two networks. The paper inherits the tokenizer and the trunk, then adds a second bank of learnable future tokens and a swappable world-model head that consumes them.
Conditional flow matching. Both heads are trained with conditional flow matching. Given a clean target and noise $\epsilon\sim\mathcal{N}(0,I)$, a scalar $\tau$ is sampled and a linear path interpolates between them; the network learns the velocity along that path. All world targets share $\tau\sim\mathrm{Beta}(1.5,1.0)$, $v$-prediction parameterization, and 50 sampling steps. Because the Beta(1.5,1.0) density puts more mass toward $\tau$ near 1, training sees near-clean states more often than near-pure noise.
Platform and data capture. The robot is two upright-mounted 6-DoF ARX5 arms with parallel-jaw grippers, a head-mounted Project Aria Gen-1 headset providing the egocentric RGB stream shared with human demonstrations, and two wrist-mounted Intel RealSense D405 cameras. Teleoperation runs through the RAIL Lab Oculus Reader interface driven by a Meta Quest 3; commanded base-frame end-effector poses are converted to joint angles by the Mink IK solver and tracked by the ARX5 joint-space controller, with low-level hardware communication over a CAN bus. Human data is captured by people wearing Aria glasses doing the task naturally. Aria Machine Perception Services recovers 21 3D hand keypoints and a 6-DoF palm pose per hand, plus a visual-inertial SLAM calibrated 6-DoF head pose, which is used both to align the action channel and to stabilize the 3D-flow target into the camera frame at time $t$.
Method
Step 1: Align the action channel as far as it goes, so BC is a strong baseline
The method section does not start with the world model. It starts by exhausting every alignment the action channel can offer, so that BC co-training is the strongest baseline the system supports rather than a strawman. Only then are WAM gains attributable to the second channel. Two datasets are in play: an egocentric human set $\mathcal{D}_{H}=\{(o^{H}_{t},a^{H}_{t})\}_{t=1}^{N_{H}}$ collected with Project Aria glasses, and a bimanual teleoperated robot set $\mathcal{D}_{R}=\{(o^{R}_{t},a^{R}_{t})\}_{t=1}^{N_{R}}$. Both share an egocentric RGB stream $I^{\text{ego}}$; the robot additionally has wrist views $I^{\text{wrist}}$.
Cross-embodiment actions are unified into a 14-D end-effector space: per-arm 6-DoF $\mathrm{SE}(3)$ pose plus a 1-D gripper command, doubled for two arms. Robot actions $a^{R}_{t:t+k}\in\mathbb{R}^{k\times d_{a}}$ come from joint angles via forward kinematics, re-expressed in the static ego-camera frame. Human hand poses $p^{H}_{t+i}\in\mathrm{SE}(3)$ are natively expressed in the moving device frame with transform $T^{\text{device}}_{t}$, so they are re-expressed in the instantaneous device frame at time $t$:
$$a^{H}_{t:t+k}=\big[(T^{\text{device}}_{t})^{-1}\,T^{\text{device}}_{t+i}\,p^{H}_{t+i}\big]_{i=1}^{k}$$
This factors out head ego-motion, so the same physical motion yields a numerically comparable trajectory across embodiments. Two residual mismatches remain. On speed, humans act faster than teleoperated robots, so embodiment-specific windows span comparable progress, $T_{H}=1\,\text{s}$ and $T_{R}=1.5\,\text{s}$, both discretized into $k$ steps; $k$ is the resampled chunk length while $T\in\{T_{H},T_{R}\}$ is the original-time horizon, and world targets $s_{t+T}$ are indexed in original time. On workspace range, quantile normalization maps each action dimension's 1st and 99th percentiles to $[-1,1]$, which is robust to hand-tracking outliers.
With aligned inputs, a shared encoder $f_{\phi}:\mathcal{O}_{H}\cup\mathcal{O}_{R}\to\mathcal{Z}$ and a shared action decoder $\pi_{\theta}(a\mid z)$ are trained end-to-end under the cross-embodiment BC objective (Eq. 1):
$$\mathcal{L}_{\text{BC-cotrain}}(\phi,\theta)=\sum_{\mathcal{D}\in\{\mathcal{D}_{H},\mathcal{D}_{R}\}}\mathbb{E}_{(o,a)\sim\mathcal{D}}\,\mathcal{L}_{\text{BC}}\!\left(\pi_{\theta}(a\mid f_{\phi}(o)),\,a\right)$$
It is realized as the sum of per-embodiment conditional flow-matching losses, $\mathcal{L}_{\text{BC-cotrain}}=\mathcal{L}_{\text{CFM}}^{\text{robot}}+\mathcal{L}_{\text{CFM}}^{\text{human}}$. The paper states plainly that this is the strongest action-aligned baseline the system supports, and that the residual transfer gap motivates the world-model interface.
Step 2: The WAM as a world-level transfer interface
To open the second channel, an auxiliary future-prediction head augments the BC policy. The shared encoder maps the current observation to a latent $z_{t}=f_{\phi}(o_{t})$, from which two parallel heads read out the modalities of interest: an action head $\pi_{\theta}(a\mid z)$ producing the chunk $a_{t:t+k}\in\mathbb{R}^{k\times d_{a}}$, and a world-model head $g_{\psi}(s\mid z)$ producing a future state $s_{t+T}$ at horizon $T$, fused within the trunk (Eq. 2):
$$p_{\theta,\psi}\big(a_{t:t+k},\,s_{t+T}\,\big|\,o_{t}\big)=p_{\psi}\big(s_{t+T}\mid z_{t}\big)\,p_{\theta}\big(a_{t:t+k}\mid z_{t}\big),\qquad z_{t}=f_{\phi}(o_{t})$$
Both heads train jointly under $\mathcal{L}_{\text{WAM}}=\mathcal{L}_{\text{action}}(a_{t:t+k})+\lambda\,\mathcal{L}_{\text{world}}(s_{t+T})$, where $\lambda$ trades action fidelity against future-prediction supervision. The key property is that $\mathcal{L}_{\text{world}}$ supervises the shared trunk through dynamics produced by both embodiments, while $\mathcal{L}_{\text{action}}$ supervises through the aligned actions already exhausted in Step 1. Human data can therefore shape the shared representation $z_{t}$ through future-scene prediction even when its action labels do not transfer faithfully.
Step 3: A single backbone and a swappable head, holding every other variable fixed
Half the value of EgoWAM is architectural. Every variant shares one heterogeneous tokenizer stack. The ego-vision stem encodes $I^{\text{ego}}_{t}$ with a ResNet-18 (output dim 256) or a pretrained DINO encoder; robot batches add a matching wrist-vision stem; a per-embodiment proprioception MLP maps the 14-D end-effector pose to 256. Each stem compresses through a query-attention block (16 latent queries, 8 heads, head dim 64). A single trunk (256 embed dim, 16 blocks, 8 heads, stochastic depth 0.1, learned domain embeddings, one-frame observation history) processes the observation tokens together with 64 learnable action tokens and 16 future tokens.
The action head $\pi_{\theta}$ is a 6-block CrossTransformer (hidden width 128, 4 heads) trained with conditional flow matching. Given a clean chunk $a_{t:t+k}$ and noise $\epsilon\sim\mathcal{N}(0,I)$, it draws $\tau\sim\text{Beta}(1.5,1.0)$ and forms
$$a^{\tau}_{t:t+k}=(1-\tau)\,\epsilon+\tau\,a_{t:t+k}$$
to initialize the action tokens, with $\tau$ embedded along the hidden dimension. Alternating self- and cross-attention blocks denoise the tokens while injecting trunk context, and a final linear layer projects them to the action dimension. The world-model head $g_{\psi}$ is conditioned on trunk embeddings and predicts $s_{t+T}$ at the embodiment-specific horizon, and the identity of that $s$ is the only independent variable the study moves.

Figure 2: Model architecture. (a) EgoWAM on an HPT backbone with modality-specific stems (ego vision, proprioception, wrist vision), learned action and future tokens, a flow-matching action head, and a swappable world-model head supplying the dynamics supervision that carries human data where actions cannot. The head supports three world targets: (b) a VAE head decoding pixel latents, (c) an RAE DINO-feature head, and (d) a 3D-flow head denoising camera-stabilized motion flow.
Step 4: Three instantiations of the world target
Pixel VAE, the reconstruction baseline. The target is a future ego frame in the latent space of a pretrained video VAE, $s=\mathrm{VAE}(I^{\text{ego}}_{t+T})$: the future frame is resized to $128\times128$ and encoded by a frozen Wan video VAE into a $16\times16\times16$ latent. The head is a diffusion transformer in two instantiations. Pixel trains a lightweight DiT from scratch (6 blocks, hidden 384, 6 heads, stride-2 patchify to 64 tokens, conditioned on the mean-pooled trunk embedding). Pixel-PT initializes from the pretrained Wan 1.3B transformer following the VACE-1.3B architecture without text conditioning (30 layers, hidden 1536, 12 heads, FFN 8960). Both minimize noise prediction:
$$\mathcal{L}^{\text{VAE}}_{\text{world}}=\mathbb{E}\big\|\epsilon-\epsilon_{\psi}(s^{\tau},\tau,f_{\phi}(o))\big\|^{2}$$
The paper's verdict is blunt: this target violates all three desiderata and serves only as the reconstruction-level baseline.
DINO features, semantic abstraction. The target becomes the DINO patch features of the future ego frame, $s=\mathrm{DINO}(I^{\text{ego}}_{t+T})$: a $16\times16$ token grid in $\mathbb{R}^{768}$ from a frozen DINOv2-B, dropping [CLS] and register tokens and applying per-token layer normalization before the loss, following RAE. The head is the RAE $\text{DiT}^{\text{DH}}$ design, pairing a 6-block 384-d DiT backbone with a shallow but wide 2-block 2048-d DDT head. The width is not cosmetic. RAE finds that diffusion in a semantic latent space requires denoiser width at least the token dimensionality (768 here), below which the flow-matching loss provably fails to converge. The objective remains noise prediction, $\mathcal{L}^{\text{RAE}}_{\text{world}}=\mathbb{E}\big\|\epsilon-\epsilon_{\psi}(s^{\tau},\tau,f_{\phi}(o))\big\|^{2}$, sampled with 50 fixed Euler steps. DINO's semantic prior abstracts appearance and aligns predictions across embodiments, but the features stay spatially indexed in image coordinates, so head ego-motion (D3) is only partially mitigated. A detail worth flagging: the trained pixel decoder is never used during EgoWAM training, because supervision happens directly in feature space and the head is discarded at inference; it is invoked only for visualization.
3D flow, spatial abstraction. The target is a dense 3D motion field over $[t,t+T]$, expressed in the camera-stabilized frame at time $t$: $s=F_{[t,t+T]}$. Raw 3D point displacement on egocentric video is dominated by ego-motion, since stationary objects induce large apparent flow whenever the wearer turns their head. The fix is to feed the pretrained dense 3D point tracker Track4World (RGB-only, predicting per-pixel 3D scene flow with metric depth, intrinsics, and camera poses) together with Aria VIO head poses, so the returned positions $X_{t},X_{t+T}$ share a consistent world frame. The future position is then mapped back to the camera frame at $t$:
$$\tilde{X}_{t+T}=(T^{\text{cam}}_{t})^{-1}\,T^{\text{cam}}_{t+T}\,X_{t+T},\qquad s=F_{[t,t+T]}=\tilde{X}_{t+T}-X_{t}$$
After stabilization a static background yields near-zero flow while manipulated objects retain motion proportional to physical displacement, abstracting dynamics from both appearance and viewpoint. For robot clips the head camera is fixed, so the transform is the identity. In implementation the flow is read on a fixed $28\times40$ (1120-point) pixel-anchor grid; to suppress tracking noise, anchors whose displacement falls below a movement threshold are discarded (2 mm for robot tracks, 10 mm for human tracks, which carry residual head motion), and the first and last 20 frames of each human clip are skipped. The head is a flow-matching decoder (4 blocks, hidden 256, 4 heads) that additionally conditions on query points $q$ uniformly sampled from the current ego frame and regresses the 3D displacement of all 1120 anchors over the 100-step horizon (target shape $100\times1120\times3$, no subsampling):
$$\mathcal{L}^{\text{Flow}}_{\text{world}}=\mathbb{E}\big\|u_{\psi}(s^{\tau},\tau,f_{\phi}(o),q)-(s_{q}-\epsilon_{q})\big\|^{2}$$
All three targets share the linear flow path $s^{\tau}=(1-\tau)\epsilon+\tau s$, $\tau\in[0,1]$, $\epsilon\sim\mathcal{N}(0,I)$, and differ only in the definition of $s_{t+T}$ and the head that consumes the trunk embeddings. The 3D-flow target satisfies all three desiderata by construction and is the most spatially grounded of the three.
Step 5: Joint training, and action-only inference
The full EgoWAM objective is trained end to end (Eq. 3, with $\lambda=1$):
$$\mathcal{L}_{\text{EgoWAM}}=\underbrace{\mathcal{L}^{\text{robot}}_{\text{action}}+\mathcal{L}^{\text{human}}_{\text{action}}}_{\mathcal{L}_{\text{action}}}+\lambda\underbrace{\big(\mathcal{L}^{\text{robot}}_{\text{world}}+\mathcal{L}^{\text{human}}_{\text{world}}\big)}_{\mathcal{L}_{\text{world}}}$$
Each step draws a mini-batch from $\mathcal{D}_{R}$ and one from $\mathcal{D}_{H}$, 32 samples each. Both pass through the shared tokenizers and trunk, after which the action head yields $\mathcal{L}_{\text{action}}$ and the world-model head yields $\mathcal{L}_{\text{world}}$. Both losses supervise the same $f_{\phi}$: when human action labels transfer weakly, world prediction compensates and shapes the trunk representation. Optimization uses AdamW (learning rate $1\times10^{-4}$, weight decay $1\times10^{-4}$) under cosine annealing ($T_{\max}=1400$, $\eta_{\min}=1\times10^{-5}$) in bf16. Every variant except Pixel-PT trains on a single NVIDIA L40S for 2000 epochs of 100 steps; Pixel-PT needs two L40S in data parallel for 1000 epochs of 100 steps to fit the 1.3B-parameter pretrained backbone in memory. Either way, end-to-end training takes roughly two days per task per method. Images are ImageNet-normalized at train and test time, with color jitter ($\pm0.1$ brightness, contrast, saturation, $\pm0.05$ hue) applied during training only.
At deployment the world-model head $g_{\psi}$ is frozen and detached from the computation graph: only the shared trunk $f_{\phi}$ and the flow-matching action head $\pi_{\theta}$ are unrolled to decode $a_{t:t+k}$ from a one-frame observation. The full policy runs on a single RTX 4090 at 30 Hz, matching the control rate of the ARX5 platform. This choice carries methodological weight. Discarding the world-model head keeps the deployment footprint identical across all four world-representation variants, so every rollout difference reported in Section 5 reflects what each target taught the trunk during training rather than test-time compute. The paper positions EgoWAM accordingly: an instrument for studying which world representation transfers, with findings deployable at BC cost.
The full data path
flowchart TB
subgraph OBS["Observation at time t"]
EGO["Ego RGB I_ego<br/>Aria Gen-1, shared by human and robot"]
WRIST["Wrist RGB I_wrist<br/>robot batches only, 2x RealSense D405"]
PROP["Proprioception<br/>14-D end-effector pose"]
end
EGO --> S1["Ego vision stem<br/>ResNet-18 or DINO, dim 256"]
WRIST --> S2["Wrist vision stem<br/>matched, dim 256"]
PROP --> S3["Proprio stem<br/>per-embodiment MLP 14 to 256"]
S1 --> TRUNK["Shared HPT trunk<br/>256 dim, 16 blocks, 8 heads<br/>domain embedding, drop path 0.1"]
S2 --> TRUNK
S3 --> TRUNK
AT["64 learnable action tokens"] --> TRUNK
FT["16 learnable future tokens<br/>WAM only"] --> TRUNK
TRUNK --> AH["Flow-matching action head<br/>CrossTransformer 6 blocks, 128 dim<br/>tau ~ Beta(1.5,1.0), 50 v-pred steps"]
AH --> ACT["Action chunk a_t:t+k<br/>k=100, 14-D per arm SE(3) + gripper"]
TRUNK --> WH["Swappable world-model head g_psi<br/>supervision only"]
WH --> W1["Pixel: Wan VAE latent 16x16x16<br/>DiT from scratch or VACE-1.3B"]
WH --> W2["DINO: DINOv2-B 16x16x768<br/>RAE DiT-DH, wide 2048 head"]
WH --> W3["3D Flow: 1120 anchors x 100 x 3<br/>Track4World + Aria VIO stabilized"]
W1 --> LOSS["L_EgoWAM = L_action + lambda * L_world<br/>lambda = 1, both streams shape f_phi"]
W2 --> LOSS
W3 --> LOSS
ACT --> INF["Deployment: world head detached<br/>trunk + action head only, RTX 4090 at 30 Hz"]
Experiments
Setup. Three bimanual tasks from the EgoVerse flagship set cover precise rigid-object manipulation, deformable manipulation, and long-horizon sequencing. In cup-on-saucer the robot must reorient a cup from a randomized initial pose and place it upright on a saucer at a randomized position, with sub-task credit of 3 points for rotating and picking up the cup, a successful handover between the two arms, and upright placement. In fold-clothes the robot three-folds a T-shirt from random configurations, with 3 points for the bottom-sleeves fold, the top-sleeves fold, and the final fold in half; success requires all three stages completed cleanly. In bag-grocery the robot opens a grocery bag and loads three items from randomized positions, worth up to 4 points (1 for opening, 1 per item), where an insertion counts only if objects are placed inside in left-to-right order.
Data. 300 to 360 teleoperated robot demonstrations per task, with randomized placements, orientations, and 4-8 object combinations. Human data spans two regimes that vary in scale and observation alignment: In-Domain human data at 1:1 with the robot data, using the same scenes and objects but unmatched viewpoint and behavior; and EgoVerse at roughly 10:1, the full EgoVerse-A flagship split per task, with diverse scenes, objects, and demonstrators and no deliberate alignment to the robot data.
Protocol. Each method is evaluated on ID (20 rollouts: seen objects and scene, randomized positions and orientations) and OOD (20 rollouts: 10 with unseen objects in the training scene, 10 with seen objects in novel scenes with varied backgrounds and table heights), for 1800 real-world rollouts in total across all methods and tasks. The reported quantities are a normalized sub-task score aggregated across rollouts and a binary success rate, with error bars as 95% finite-sample-valid confidence intervals carrying Type-I error control for miscoverage. Initial conditions are randomized but held common across methods within each task.

Figure 3: Quantitative comparison on real-world rollouts. Normalized score and success rate across three bimanual tasks under ID and OOD evaluation, comparing BC against four WAM variants. WAM co-training consistently outperforms BC: human data often degrades BC yet yields large WAM gains. Pixel transfers weakly, DINO drives the strongest OOD generalization, and 3D flow the largest ID spatial gains.
Q1: WAM versus BC co-training under natural human data
When robot and human action patterns are not well aligned, for instance a human hand holding a cup sideways where the gripper must grasp it from the front, BC degrades by reproducing those inexecutable motions, while WAM consistently turns the same data into reliable gains. The qualitative comparison in Figure 5 makes the failure modes concrete: BC produces human-like motion while Pixel fails at precise cup transport; BC cannot grasp unseen objects while Pixel hallucinates a free-space grasp; BC overfits and fails to reach the cloth while Pixel's geometric confusion leaves the third fold incomplete. Even when large amounts of human data are introduced, BC tends to overfit the robot data alone and fails to benefit from human data for generalization. On fold-clothes, BC + EgoVerse still overfits and cannot adapt to novel scenes with lower table heights, whereas the DINO-based WAM leverages the in-the-wild human data to generalize to new objects and scenes.
The paper traces this to the representation itself. A UMAP of trunk embeddings on cup-on-saucer shows BC isolating human and robot embeddings (and even human from human) under unaligned action supervision, whereas WAM aligns them into a shared space through the additional task-relevant dynamics supervision. That is the mechanistic evidence behind the claim that WAM co-training is what makes 10:1 human data usable, rather than a claim resting on score curves alone.

Figure 4: UMAP of trunk embeddings on cup-on-saucer. BC separates human and robot embeddings; WAM aligns them in a shared latent.
Q2: The world representation decides what transfers
Pixel-based prediction transfers weakly, because reconstructing raw appearance entangles embodiment- and scene-specific detail that does not travel, leaving the hallucination and geometric confusion visible in Figure 5. Abstraction repairs this in two complementary ways. The DINO-based WAM predicts semantic features and yields the strongest OOD generalization to unseen objects and scenes, up to 4x in the abstract's terms. The 3D-flow WAM predicts camera-stabilized motion and delivers the largest in-domain spatial gains, succeeding at precise cup placement across the workspace as mapped in Figure 6, for a 20-30% in-domain improvement. Figure 7 grounds the contrast in the predictions themselves: pixel VAE prediction barely changes with co-training, the semantic RAE target refines object shape from human data but only moderately, and the 3D-flow target benefits most, recovering object motion that robot-only prediction leaves static.

Figure 6: Spatial gains for 3D flow. 3D flow succeeds (green) where BC fails (red) across cup positions in the workspace.
Q3: Three ablations that pin down causality
Unaligned human data (Figure 8). Of the three tasks only bag-grocery benefits from action-level co-training, because its pick-and-place motions are naturally aligned between human and robot. As a counterfactual the authors collect unaligned human demonstrations in which the demonstrator picks objects in an unusual manner that retargets to inexecutable robot actions. Under that data, BC collapses below the robot-only baseline, while the 3D-flow WAM stays robust and still surpasses robot-only.
Aligned human data (Figure 9, Appendix A.1). The opposite extreme: the human demonstrator deliberately mirrors the robot's viewpoint and grasp strategy, with camera height and viewpoint matched to the robot's static Aria mount, on cup-on-saucer, the task with the strongest human-robot mismatch, under two 1:1 human regimes (In-Domain Co-train with natural human data, and Aligned Co-train). Three findings come out. Aligned human data lifts BC above its robot-only baseline, confirming that BC benefits from human data only when the action distribution is hand-curated to match the robot's. Pixel and DINO improve when ego-motion is factored out at collection time, Pixel rising from 35% to 65% and DINO from 50% to 70%, and that 20-30 point gap quantifies how much image-coordinate targets suffer when (D3) is violated. The 3D-flow WAM is invariant to alignment, holding at 85% under both regimes, because the camera-stabilized target factors out ego-motion by construction and therefore gets for free what the others recover through manual alignment.
Human data modality (Figure 11, Appendix A.2). How much of the gain comes from action labels on human demonstrations versus world-model supervision? Three co-training modalities are ablated on cup-on-saucer: Action Only (the BC co-train baseline), 3D Flow only (no action supervision on human batches), and Action + 3D Flow (full EgoWAM). World-model supervision alone outperforms action supervision alone: 3D-flow-only beats action-only across all three splits, most starkly on OOD Scene, where action-only fails entirely at 0% success rate while 3D-flow-only retains 10%. The two are also mutually reinforcing, since joint training wins on every split. The paper attributes this to two complementary effects. Action as context: action labels condition the trunk on the demonstrator's intent and sharpen the world-model prediction, measurable directly in the world-model loss. Action as a task-relevance signal: action labels mark which motion in the scene is task-relevant, focusing the trunk on agent-caused dynamics rather than incidental flow.
| 3D-flow prediction stream (cup-on-saucer) | Flow-Only | Action + Flow (full EgoWAM) |
|---|---|---|
| Human stream | 0.23 | 0.22 |
| Robot stream | 0.20 | 0.19 |
Table 1 (paper Table 1): 3D-flow world-model prediction loss. Adding action labels lowers the loss on both the human and the robot stream, which is the direct evidence for action-as-context.
| Task | Robot demos / hours | Human In-Domain (hours) | Human EgoVerse (hours) |
|---|---|---|---|
| cup-on-saucer | 300 / 2.5 h | 2 h | 20.5 h |
| fold-clothes | 360 / 3.0 h | 2 h | 21 h |
| bag-grocery | 300 / 2.5 h | 2 h | 7 h |
Table 2 (paper Table 3): Data composition per task. In-Domain is small-scale human data aligned in scene and object at a 1:1 ratio with robot demos; EgoVerse is the large-scale EgoVerse-A flagship split, giving an overall human-to-robot ratio near 10:1.
Simulation replication: robot-to-robot transfer on RoboTwin
To validate the same two claims in a public, reproducible robot-to-robot setting, EgoWAM is instantiated on the RoboTwin 2.0 bimanual benchmark running on SAPIEN. The primary embodiment is aloha-agilex with two 6-DoF arms, co-trained with arx-x5, franka, and ur5, the latter two being 7-DoF, on three hard tasks: pick-diverse-bottles with a 15-instance bottle pool, stack-bowls-three as long-horizon manipulation, and hanging-mug as fine-grained manipulation, all using the shipped demo_clean-50 split. Because the arms differ in degrees of freedom, both arms predict an absolute end-effector action in a head-camera frame, $[\mathbf{p}_{xyz},\boldsymbol{\theta}_{\text{ZYX}},g]$ per arm for a 14-D vector identical across robots, so a single action head serves every embodiment. At test time the predicted chunk is resampled to the control rate and streamed to a differential-IK controller under a receding horizon. Each world-model variant is trained single (aloha-agilex only) and cross (all four robots) per task, holding trunk, head, and optimizer fixed. As same-action-space references, RoboTwin's official ACT and Diffusion Policy pipelines are adapted into ACT-EE and DP-EE by replacing their native joint-space state and action with the same 14-D end-effector vector and evaluating through the identical IK executor; they are single-embodiment and world-model-free, which isolates what cross-embodiment co-training and the world head add on top of the action space alone. All methods are evaluated on 100 held-out seeds.
| Task | ACT-EE | DP-EE | BC s | BC c | Pixel s | Pixel c | DINO s | DINO c | 3D Flow s | 3D Flow c |
|---|---|---|---|---|---|---|---|---|---|---|
| pick-diverse-bottles | 2 | 5 | 2 | 6 | 7 | 11 | 4 | 28 | 0 | 16 |
| stack-bowls-three † | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 8 | 0 | 16 |
| hanging-mug | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 |
Table 3 (paper Table 4): RoboTwin held-out closed-loop success in percent over 100 seeds. s = single-embodiment training, c = cross-embodiment co-training, best per task in bold. † stack-bowls-three is evaluated on an appearance-shifted object: an incidental asset change between the released demonstrations and the current simulator shifted the bowl's material while its pose and geometry stayed the same, and demo_clean adds no texture randomization, which turns the task into an appearance-shift test.
Four things carry weight in that table. Cross-embodiment co-training beats single-embodiment training: every variant improves from single to cross (on pick, DINO 4 to 28, Pixel 7 to 11, BC 2 to 6), and cross is the only setting that succeeds at all on stack. Since single-embodiment policies already overfit their training scenes yet barely generalize, the gap is genuine cross-embodiment transfer rather than better fitting. The world-model head, not the action space, drives the transfer: under cross-embodiment training the world-model variants consistently exceed the action-only BC policy, while the same-action-space but world-model-free references ACT-EE and DP-EE stay at or below the single-embodiment variants (at most 5% on pick, 0% elsewhere), so the gain is attributable to co-training through the world-model interface. Appearance abstraction (D1) is directly reflected: on stack, only the appearance-invariant DINO at 8% and 3D flow at 16% survive the material shift, while BC and Pixel collapse to 0%. And very fine-grained manipulation still fails for everyone: hanging-mug requires threading a mug onto a thin rack with millimeter precision, and every method including DINO and 3D flow stays at or below 1%, meaning the bottleneck there is manipulation precision rather than object generalization or the world target.
A telling failure mode: a pretrained pixel model hallucinates that the task is already done
On bag-grocery, Pixel-PT trained on robot data only underperforms both BC and Pixel. The failures concentrate at the bag-opening stage, and Figure 12 isolates the cause. Pixel-PT renders crisp future frames in which the bag already appears open before the gripper has acted on the handles: the natural-image prior, in which bags are typically open, overrides the actual scene state, and the policy advances to the pick stage prematurely. Pixel trained from scratch produces blurrier frames but tracks bag openness faithfully, and the resulting policy opens the bag more reliably. Pixel-PT co-trained with EgoVerse keeps the sharp pretrained output while predicting the closed-then-opening progression correctly, which indicates that scaling up data through human co-training can correct hallucination induced by a noisy prior in a pretrained video model.

Figure 12: Video prediction for failure analysis on bag-grocery, six-step rollouts from the same initial frame. Pixel with robot only: blurry but faithful, the bag stays closed until acted on. Pixel-PT with robot only: hallucinates an already-open bag, causing the policy to skip the opening stage. Pixel-PT with EgoVerse: sharp and faithful, human co-training removes the hallucination.
Limitations
The authors state three limitations, and all three are worth recording as written. First, motion generalization: the gains stay at the context level, and learning a novel motion primitive or skill from human data, for example folding shorts from a T-shirt model, remains out of reach; a more unified action representation is left to future study. This is consistent with the RoboTwin result where hanging-mug is at zero for every method, since representation alignment addresses scene understanding rather than manual dexterity. Second, multi-task scaling: the study isolates the world representation with one policy per task, and multi-task co-training on large-scale in-the-wild human data remains to be explored. Third, the world representation is still open: the paper describes itself as a starting point showing that DINO and 3D flow beat pixels for cross-embodiment transfer, while the best world representation for scaling robot learning remains a valuable open research question.
Beyond the authors' own list, two further constraints are visible in the numbers. First, evaluation sample sizes are small. Each method gets 20 rollouts per task per split, and the 1800 figure is the total across all methods, tasks, and ID/OOD conditions. The paper uses 95% finite-sample-valid confidence intervals with miscoverage control, but at success rates in the single digits to a few tens of percent, 20 rollouts resolve in steps of 5 percentage points, so orderings such as BC versus Pixel on bag-grocery are not obviously stable. The simulation side is steadier at 100 seeds, yet absolute success rates are so low (28% at best on pick, mostly zeros elsewhere) that all three RoboTwin tasks remain too hard for every method to separate DINO from 3D flow cleanly. Second, the three targets are not equally cheap or equally self-contained. The 3D-flow target depends on the Track4World dense tracker plus Aria VIO head poses, DINO depends on a frozen DINOv2-B and RAE's wide-head design, and Pixel-PT depends on a 1.3B-parameter VACE backbone that forces two GPUs. The controlled part of the study is the trunk, action head, and data mixture; the external pretrained models and labeling pipelines behind each target are not controlled, so the paper does not separate how much of the 3D-flow advantage comes from the geometric construction that factors out ego-motion and how much comes from Track4World's own priors. The 3D-flow target is also a dense $100\times1120\times3$ tensor with no subsampling, so the supervision dimensionality is on a different scale from the other two targets, and a fixed $\lambda=1$ cannot have an equivalent effective weight across them; no sensitivity analysis over $\lambda$ is reported.
Conclusion and Outlook
The contribution separates into three layers. As a framework, EgoWAM provides a single backbone with a swappable world-model head, enabling apples-to-apples comparison of world representations under identical action supervision and data mixtures, while discarding the world-model head at inference so deployment costs what BC costs. As evidence, it shows that WAM co-training unlocks human-data scale: across in-domain and OOD evaluations it scales and transfers more consistently from large-scale natural egocentric human data than BC co-training, confirming that future-dynamics supervision is the missing channel where action-only supervision saturates. As direction, it establishes the world representation as the next critical axis, with DINO and 3D flow both substantially outperforming the pixel-VAE baseline and complementary strengths, DINO's semantic prior giving the strongest object and scene generalization and 3D flow's geometric grounding giving the strongest spatial generalization and the highest overall scores.
For teams building robot learning systems, three results are directly actionable. If the human data has not been aligned in viewpoint and behavior, do not expect action-level co-training to pay; it is more likely to cost performance. Attaching a future-prediction head that exists only during training is a cheap intervention, because it leaves the inference path and test-time compute untouched while changing what the backbone learns. And when choosing what to predict, screen candidates against appearance abstraction, cross-embodiment consistency, and ego-motion factoring: image-coordinate targets, pixel VAE and DINO features alike, pay 20 to 30 points when the human head motion at collection time is uncontrolled, whereas defining the target on camera-stabilized 3D motion removes that cost by construction. The authors call this set of desiderata a recipe for further study and position human data as a flywheel for scalable robot learning, explicitly offering the work not as a definitive answer but as a point of departure that turns "use human data" into a concrete design axis.
Golden Quotes
"The shared action decoder is the sole channel through which human data reaches the policy, forcing it to entangle transferable content (objects, scenes, semantics) with non-transferable execution (morphology, behavioral style), and the embodiment gap in the latter blocks transfer of the former." Section 1 frames the failure of action-level co-training as an interface problem, not a data-volume problem.
"Targets should represent the effect rather than the agent, so a human hand and a robot gripper producing similar physical change induce similar supervision." Desideratum (D2), and the most compact answer the paper gives to the question of what to predict.
"Discarding the world-model head at inference also keeps the deployment footprint identical across all four world-representation variants, so any rollout differences in Sec. 5 reflect what each target taught the trunk during training rather than test-time compute." Appendix B.3, the methodological statement that makes this readable as a controlled experiment.



