PAPER DEEP DIVE
MachEmbodied-U0: Unified Understanding and Generation Model for Embodied Intelligence
General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate precise actions. Vision-language-action models provide strong semantic priors but typically do not explicitly model scene dynamics, while world-action models couple visual prediction with control without necessarily exposing the task-relevant semantic and spatial structure needed for fine-grained manipulation. We present MachEmbodied-U0 (ME-U0), a unified embodied foundation model connecting understanding and generation experts through a Mixture-of-Transformers architecture. Subtask prediction and affordance grounding guide joint visual-dynamics and action generation via flow matching. Visual dynamics encompass future RGB, depth, surface normals, and optical flow, providing complementary supervision for appearance, geometry, and motion. Multi-rate Rotary Position Encoding (MRPE) aligns visual dynamics with fine-grained control. We pretrain ME-U0 on approximately 4,200 hours of curated demonstrations from robotic datasets and egocentric datasets. Using only the supervision natively available in each downstream benchmark, ME-U0 achieves an average score of 17.66 on the RoboDojo simulation benchmark and average success rates of 99.0% and 82.5% on LIBERO and LIBERO-Plus, respectively. We additionally validate ME-U0 on real-world robotic manipulation tasks, demonstrating its effectiveness beyond simulation. Without corresponding downstream supervision, ME-U0 further demonstrates zero-shot subtask prediction, affordance grounding, and visual dynamics on simulated and real-world observations. Overall, ME-U0 combines competitive downstream control performance with transferable task-grounding and visual-dynamics capabilities across simulation and the real world.
Source: arXiv:2609.25627 (cs.RO), Foundation Model team, Li Auto Inc., submitted Sep 22, 2026. Project page: machembodied.com/ME-U/ME-U0.html; code: github.com/MachEmbodied/ME-U0.
In One Sentence
ME-U0 puts an "understanding" expert and a "generation" expert inside a single Mixture-of-Transformers backbone: the understanding expert emits subtask descriptions and affordance grounding, while the generation expert jointly denoises future visual dynamics (RGB / depth / surface normals / optical flow) together with continuous robot actions via flow matching. Pretrained on roughly 4,200 hours of curated robot + egocentric data, it scores 17.66 on RoboDojo, 99.0% on LIBERO and 82.5% on LIBERO-Plus, and retains zero-shot subtask, affordance and visual-dynamics capabilities even without any downstream supervision for them.
1. The Problem It Tackles
General-purpose robot control needs four things at once: understanding task intent (what to do), knowing where to interact (where to act), capturing how the scene evolves (how the scene evolves), and generating precise actions (how to realize it).
Each existing line covers only half:
- VLA models (RT-1 / RT-2 / Octo / OpenVLA / $\pi_0$ / $\pi_{0.5}$) carry strong semantic priors, but their action-centric objectives typically do not explicitly model how the scene evolves during interaction.
- WAMs (UniPi / GR-1 / DreamZero / LaWAM / Fast-WAM) couple future-observation prediction with action learning, providing dense spatial and motion supervision — but visual dynamics alone does not ensure task-relevant grounding, interaction localization, or long-horizon task decomposition.
ME-U0's stance: these should not be an either-or choice; inside one architecture, each should condition the other. Semantic understanding supplies context for generation; visual generation feeds anticipated physical consequences back into control.
Task: Twist the bottle cap open"] --> U["Understanding Expert"] V["Multi-view observation + robot state"] --> U V --> G["Generation Expert"] U -->|"Subtask + Affordance
(shared multimodal attention)"| G G -->|"Joint flow-matching denoising"| D["Visual dynamics
RGB / Depth / Normal / Flow"] G -->|"Joint flow-matching denoising"| Act["Continuous actions
Action chunk $a_{t:t+H_a}$"] Act --> R["Robot execution"] D -.->|"Physical-consequence supervision"| Act
Figure 1: ME-U0 overview — cooperation between the understanding and generation experts.
2. Architecture: A MoT with Two Experts
2.1 Overall Backbone
The model is initialized from Lance (an existing unified multimodal understanding + generation model) and further trained for robot capabilities. The sequence layout is: text tokens → current-observation and robot-state tokens → future visual and action tokens (noisy). Causal attention prevents text tokens from attending to future targets; during denoising the generation expert reaches the understanding expert's representations of input images and text through shared multimodal attention.
Each expert owns its FFN parameters (the core of MoT: per-expert parameters, shared attention), so the very different demands of semantic understanding and continuous generation do not drag on each other, while still coordinating through shared attention.
2.2 Understanding Expert: Subtasks + Affordances
Both abilities are formulated as autoregressive text generation trained with next-token cross-entropy:
$$\mathcal{L}_{\text{und}} = -\frac{1}{T}\sum_{t=1}^{T}\log p_\theta\!\left(y_t \mid y_{<t}, \mathbf{x}\right)$$
where $\mathbf{x}$ contains the current image, task instruction and contextual metadata, $y=(y_1,\dots,y_T)$ is the target response containing subtask and affordance labels, and the loss is computed only on supervised response tokens.
- Subtask prediction: the overall instruction may stay unchanged across manipulation stages while the immediate goal shifts with the scene. The model must describe the local manipulation goal given the instruction and the current observation — following the semantic subtask supervision of $\pi_{0.5}$.
- Affordance prediction: for annotated keyframes, the same QA formulation asks the model to identify the task-relevant target in the current image, predict its bounding box, then the intended interaction and interaction point. When both labels exist, affordance follows subtask in the same assistant response; samples without affordance annotations keep their original objective.
Figure 6: Framework — the understanding expert predicts subtasks and affordances while the generation expert jointly predicts future visuals and actions via shared multimodal attention.
2.3 Generation Expert: Joint Visual-Dynamics + Action Denoising
Given conditioning context $c$, clean visual targets $z$ and action trajectories $a$ are corrupted with independent Gaussian noise at a shared flow time $\tau$:
$$\tau = \sigma(u + \log 4),\quad u \sim \mathcal{N}(0,1)$$
$$z_\tau = (1-\tau)z + \tau\boldsymbol{\epsilon}_z,\qquad a_\tau = (1-\tau)a + \tau\boldsymbol{\epsilon}_a$$
This is a shifted logit-normal distribution, deliberately weighting sampling toward higher noise levels. The noisy visual and action tokens go together into the generation expert, exchange information via bidirectional self-attention, and separate output projections predict the velocity fields:
$$(\hat{u}_z, \hat{u}_a) = v_\theta(z_\tau, a_\tau, \tau; c)$$
The joint flow-matching objective is:
$$\mathcal{L}_{\text{gen}} = \mathbb{E}\left[\lambda_z\left\|\hat{u}_z - (\boldsymbol{\epsilon}_z - z)\right\|^2_{M_z} + \lambda_a\left\|\hat{u}_a - (\boldsymbol{\epsilon}_a - a)\right\|^2_{M_a}\right]$$
with $\|\cdot\|^2_M$ the mean squared error over elements selected by mask $M$. The total objective combines both:
$$\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{gen}} + \lambda_{\text{und}}\mathcal{L}_{\text{und}}$$
Why insist on joint denoising instead of a separate action denoiser? The paper is direct: a separate action denoiser would decouple action generation from the shared generative process and weaken the architectural unification of understanding, visual prediction, and robot control. Joint denoising lets actions participate in multimodal generation and interact with visual predictions throughout denoising.
The same backbone supports forward and inverse dynamics by varying conditioning and targets:
$$\text{Forward: } (c, a) \rightarrow z,\qquad \text{Inverse: } (c, z) \rightarrow a,\qquad \text{Joint: } c \rightarrow (z, a)$$
2.4 Multimodal Visual Targets: Not Just RGB
Visual dynamics covers four modalities: future RGB, depth, surface normals, and optical flow. Depth and normals are represented as three-channel visual targets, as is flow — supplying complementary supervision for geometry, surface orientation, and inter-frame motion. Crucially, all modalities share one video VAE latent space and one generation backbone; no modality-specific prediction heads are needed. The prompt's video-modality field selects which one to generate.
2.5 MRPE: Aligning Actions with Video on the Time Axis
An easy-to-miss but engineering-critical detail. The video VAE compresses time, and actions are sampled three times as densely as video — each video latent frame corresponds to a whole chunk of consecutive actions. If actions within a chunk shared one position index, their internal order would be inexpressible.
MRPE (Multi-rate Rotary Position Encoding) does something tiny on top of mRoPE: keep the shared temporal position of the corresponding video latent frame, but assign each action within the chunk a distinct height-axis index while keeping the width-axis index fixed. Temporal alignment between actions and video is preserved while each action's position within its chunk is explicitly encoded. The sampling rule ensures:
$$H_a = 3(H_v - 1)$$
i.e., $H_a$ action steps and $H_v$ video frames (including the anchor), exactly three actions between consecutive sampled video frames; after the video VAE's temporal compression this yields a consistent 12 action steps per predicted video-latent step across datasets.
2.6 Unified Action Representation: Six Embodiments in One Space
Embodiment diversity brings rich experience, but inconsistent control conventions obscure the common structure across demonstrations — similar values may describe different physical quantities or even opposite control directions. The paper defines a unified state/action interface:
| Component | Dim. | Shared representation |
|---|---|---|
| Arm: joint control | 6 / 7 | Joint configuration in radians |
| Arm: EEF control | 9 | Cartesian position in metres and 6D orientation |
| Gripper | 1 | Normalized closing coordinate: $0$ = open, $1$ = closed |
| Chassis | 3 | Body-frame planar velocity $(v_x, v_y, \omega_z)$ in m/s and rad/s |
| Torso: AgiBot | 2 | Pitch and lift position, radians and metres |
| Torso: Galaxea | 3 | Planar Cartesian velocity $(v_x, v_z, \omega_y)$ in m/s and rad/s |
| Head | 2 | Yaw and pitch joint positions in radians |
Concrete examples show the workload: Galaxea's native gripper commands are $[0,100]$ and must be reversed and rescaled into the closing coordinate while AgiBot's convention is kept; AgiBot's millimetre gripper states are calibrated separately. Galaxea records mobile-base states as wheel-steering angles and wheel velocities while actions are Cartesian velocities — a kinematic conversion is needed so states and actions share a physical interpretation. Quantities with genuinely distinct meanings (torso position vs torso velocity) stay explicitly separated rather than merged.
Figure 2: Robot data composition and scene–action–object diversity.
3. Data: From 5,700 + 3,920 Hours Down to 4,200
3.1 Composition
The raw pool is roughly 5,700 hours of robot demonstrations + 3,920 hours of egocentric demonstrations; after curation about 4,200 hours are retained. Robot data comes from four open-source datasets — AgiBot World Beta, RoboMIND 2.0, RoboCOIN, Galaxea Open-World — spanning six embodiments (AgiBot G1 at 52.8%, plus Galaxea R1 Lite, AgileX COBOT, Magic V2.0, ARX LIFT, Dual-Arm Franka Research 3, Realman RMC AIDA-L).
A scene–action–object tag scheme (aggregated into six scene groups, six action groups, eight object groups) verifies that shared behaviors recur across task contexts rather than covering one narrow behavior set.
Egocentric data (EgoDex, EgoLive, EgoVerse, HOI4D, HOT3D, and the hand-annotated part of Ego-Exo-4D) backfills task tags underrepresented in robot data, requiring temporally aligned first-person observations plus low-level human motion annotations (fingertip locations, wrist poses, arm/shoulder states).
3.2 Three-Stage Curation
scale each dim by $q_{99}-q_{01}$
median filter + smoothing as reference
large residual/acceleration/jerk → drop"] C1 --> C2["② State-action consistency
estimate command-response lag via cross-correlation
of smoothed first differences; check directional agreement
low agreement / state freeze while action varies → drop"] C2 --> C3["③ Physical bounds
manually reviewed per-dimension bounds
(gripper channels exempt — open/close is not smooth motion)"] C3 --> K["Keep episode
original recordings untouched, temporal continuity preserved"]
The paper gives three concrete failures: desktop wiping at 29.93 s with J6 residual 1.907 vs threshold 0.054; sliding bar stool at 9.40 s with action $+0.682$ but state $-0.657$ (opposite directions); storing tableware at 81.07 s with J1 target $-6.26$ below the bound $-2.00$.
Figure 5: Egocentric data processing — Ego2Robot synthesis converts human demonstrations into robot-aligned observations and trajectory annotations.
3.3 Egocentric → Robot: Ego2Robot Synthesis
This pipeline is the most engineering-heavy part of the paper. Episodes are first categorized by object attributes into easy (trajectory replay) and hard (contact refinement):
- Shared steps: metric depth from Depth Anything 3 + masks from SAM3 → following Inpaint Anything, ProPainter removes human regions and recovers the occluded background.
- Easy: replay annotated trajectories, using CoWTracker to ensure pixel-space accuracy and reduce 2D–3D mismatch.
- Hard: SAM3D reconstructs manipulated rigid objects as 3D physical models; masks/depth/6D poses are fused into temporally consistent object trajectories; wrist and fingertip trajectories are then refined by a two-stage PPO curriculum for contact accuracy and smoothness.
- Finally: the best trajectories from both branches are retargeted to gripper-equipped embodiments via robot-specific inverse kinematics (Mink in MuJoCo), and the robot mesh is rendered into the inpainted egocentric video with occlusions resolved by object depth.
3.4 Task-Aware Sampling
Data diversity ≠ training-exposure diversity. Sampling proportional to timesteps concentrates supervision on abundant tasks and long demonstrations. The fix is global task-aware sampling: first give every eligible task an initial coverage allocation drawn from multiple episodes, then distribute the remaining budget by temperature sampling:
$$w_k = \max(R_k, 1)^{\alpha}$$
where $R_k$ is the task's remaining capacity and $\alpha = 0.5$. Sublinear weighting reduces the dominance of large task buckets while retaining a preference for tasks with more data; allocations stay bounded by available capacity to avoid aggressive oversampling of few-demo tasks. Capacity-short tasks are supplemented from egocentric data. In total 120 million training samples are drawn.
3.5 Per-Source Video–Action Temporal Alignment
| Source | Actions (Hz) | Video (Hz) | $H_a / H_v$ | Horizon (s) |
|---|---|---|---|---|
| AgiBot World Beta | 30 | 10 | 36 / 13 | 1.20 |
| Galaxea Open World | 15 | 5 | 24 / 9 | 1.60 |
| RoboMIND 2.0: AgileX / Mobile / ARK | 30 | 10 | 36 / 13 | 1.20 |
| RoboMIND 2.0: Franka | 15 | 5 | 24 / 9 | 1.60 |
| RoboCOIN: 30 Hz recordings | 30 | 10 | 36 / 13 | 1.20 |
| RoboCOIN: 50 Hz recordings | 25 | 8.33 | 36 / 13 | 1.44 |
The constraints are explicit: prediction windows stay within ~1–2 seconds, and actions are sampled three times as densely as video — finer control supervision bought without equally dense visual sequences.
4. Training Infrastructure (Refreshingly Concrete)
Two recurring costs got dedicated optimizations:
- Dataset initialization: episode manifests are precomputed at data-preparation time (ids, lengths, task associations, filtering info); a reusable subtask index stores full annotated intervals from which anchors are derived per configuration; lightweight LeRobot readers read native formats directly. Result: full-mixture initialization drops from ~1 hour to under 10 minutes.
- Video decoding: training samples are organized into logical shards (same task, nearby anchors); workers merge overlapping frame requests — in the paper's example 12 requests collapse to 6 unique frames, decoded into shared memory for assembling multiple samples. Add bounded decode-job spans, concurrency limits, asynchronous prefetch of the next shard, and offloading decoding to separate CPU jobs writing a shared cache. Result: overall training time reduced by more than 30%.
5. Experiments
5.1 The Key Premise: No Extra Downstream Annotations
This has to be stated up front, or the numbers get misread. Neither RoboDojo nor LIBERO provides subtask/affordance annotations or auxiliary geometry/motion supervision, and for fair comparison the authors add none during post-training. In other words, the subtask, affordance, depth/normal/flow capabilities learned in pretraining sit idle — and possibly degraded — during downstream fine-tuning. The results below are reported exactly under this reduced-supervision setting.
Post-training setup: both benchmarks start from the same pretrained checkpoint, each using 368 PPU810E accelerators with per-device batch size 2; LIBERO uses an action horizon of 24 with EEF pose control, RoboDojo a horizon of 48 with delta joint-position control.
5.2 RoboDojo-Sim: Strongest Among WAMs
| Model | Class | Overall SR | Overall Score | Precision Score | Long-Horizon Score |
|---|---|---|---|---|---|
| StarVLA | VLA | 3.24 | 6.40 | 9.90 | 14.15 |
| X-VLA | VLA | 6.52 | 10.13 | 18.32 | 16.53 |
| $\pi_{0.5}$ | VLA | 6.93 | 11.44 | 12.40 | 23.54 |
| Spatial Forcing | VLA | 8.04 | 12.38 | 17.32 | 23.26 |
| Hy-Embodied-0.5-VLA | VLA | 8.80 | 13.07 | 13.81 | 25.74 |
| Xiaomi-Robotics-1 | VLA | 13.93 | 20.07 | 26.69 | 38.39 |
| GalaxeaVLA (G0.5) | VLA | 14.88 | 20.23 | 28.25 | 44.12 |
| DM0.5 | VLA | 19.34 | 24.90 | 24.82 | 33.70 |
| Fast-WAM | WAM | 2.03 | 3.48 | 1.96 | 9.14 |
| AHA-WAM | WAM | 2.39 | 4.82 | 5.86 | 8.61 |
| GigaWorld-Policy-0 | WAM | 3.27 | 6.20 | 6.15 | 15.51 |
| X-WAM | WAM | 3.83 | 7.69 | 6.72 | 17.47 |
| OpenWAM-$\alpha$ | WAM | 11.92 | 17.18 | 18.45 | 34.93 |
| ME-U0 (ours) | WAM | 11.18 | 17.66 | 23.95 | 36.98 |
Three readings:
- Strongest overall among WAMs (17.66 vs OpenWAM-$\alpha$'s 17.18), with the clearest leads in precision (23.95) and long-horizon (36.98) — matching the paper's claim that grounding joint visual–action dynamics with subtask and affordance context benefits both fine-grained interaction and extended execution.
- With only ~4,200 hours of pretraining it stays competitive with mature VLA baselines (17.66 above $\pi_{0.5}$'s 11.44 and Spatial Forcing's 12.38).
- Honest about the weakness: memory-dependent tasks are the worst dimension, Score 8.42 / SR 7.00%. The paper attributes this to the model not explicitly retaining observation history — once task-relevant information leaves the current view, it is lost. This is its self-identified next step.
5.3 LIBERO and LIBERO-Plus
| Model | Class | Spatial | Object | Goal | Long | Average |
|---|---|---|---|---|---|---|
| $\pi_{0.5}$ | VLA | 98.8 | 98.2 | 98.0 | 92.4 | 96.9 |
| OpenVLA-OFT | VLA | 97.6 | 98.4 | 97.9 | 94.5 | 97.1 |
| ABot-M0 | VLA | 98.8 | 99.8 | 99.0 | 96.6 | 98.6 |
| X-VLA | VLA | 98.2 | 98.6 | 97.8 | 97.6 | 98.1 |
| Fast-WAM | WAM | 98.2 | 100.0 | 97.0 | 95.2 | 97.6 |
| LingBot-VA | WAM | 98.5 | 99.6 | 97.2 | 98.5 | 98.5 |
| OpenWAM-$\alpha$ | WAM | 99.6 | 99.6 | 99.8 | 98.2 | 99.3 |
| ME-U0 (ours) | WAM | 98.4 | 99.8 | 99.0 | 98.6 | 99.0 |
Standard LIBERO is nearly saturated (99.0%, all four suites between 98.4–99.8%), so the informative benchmark is LIBERO-Plus — the same checkpoint, zero adaptation, on 10,030 perturbed instances: 82.5% overall:
| Perturbation | ME-U0 | Best compared |
|---|---|---|
| Light | 97.6 | ImageWAM 98.1 |
| Language | 90.4 | ImageWAM 91.4 |
| Layout | 84.5 | Qwen-RobotManip 87.3 |
| Background | 81.9 | Qwen-RobotManip 97.7 |
| Noise | 81.3 | Qwen-RobotManip 97.7 |
| Robot initial state | 75.7 | OpenWAM-$\alpha$ 76.1 |
| Camera | 70.2 | Qwen-RobotManip 87.2 |
| Overall | 82.5 | Qwen-RobotManip 89.0 / $\pi_{0.5}$ 84.4 |
The conclusion is crisp: lighting and language perturbations barely threaten it (97.6% / 90.4%), while camera-view changes and robot initial states are the two hardest categories (70.2% / 75.7%) — both explicitly named as future work.
5.4 Real-World Deployment and Zero-Shot Analysis
Real-world deployment uses the pretrained checkpoint with no post-training or deployment-specific adaptation, inference on a remote server. Three demo tasks: picking up a marker, placing a tissue into a box, and positioning the number 20 to complete $3+17=20$ (the last one is telling — it requires grounding semantics into specific object selection). The paper is candid: these platforms all appear in the pretraining corpus, so this tests deployment on familiar embodiments, not generalization to unseen robots.
Zero-shot analysis (all samples outside the pretraining corpus, used only for evaluation):
- Visual dynamics: on RoboDojo simulation and real-world splits, generated depth, normals and flow capture meaningful scene structure and motion. Real-world predictions are slightly more consistent than simulated ones — the authors conjecture a domain gap between simulated imagery and the real observations seen during pretraining.
- Affordances: on RoboDojo sim/real, self-collected data, and RDT-1B data, the model names task-relevant targets with bounding boxes — chargers, fruit, pens, straws, plus non-object regions like bowl interiors and table surfaces.
- Subtasks: on the same samples it describes the current stage, e.g., "pick up the charger on the table with right gripper" → "insert the charger into the power strip".
6. Our Take
What makes this paper worth attention is not any single metric but that it puts understanding, prediction, and control inside one trainable object — and gives a clear ablation hint: the large leads on RoboDojo's precision and long-horizon dimensions are hard to explain by "more data" alone; more likely, subtask/affordance context genuinely helps the generation expert find where to look.
Equally, its boundaries should be seen:
- RoboDojo overall SR is only 11.18% — the benchmark itself is brutally hard (42 tasks, open vocabulary); the best VLA is 19.34%. Don't extrapolate LIBERO's 99% onto it.
- Those elegant zero-shot capabilities sit idle during downstream fine-tuning; the paper never shows the gain from keeping their supervision on — a real gap.
- Memory is an acknowledged weakness (Score 8.42); the model keeps no observation history — essentially a current-frame-conditioned policy.
- 70.2% under camera perturbation on LIBERO-Plus shows limited robustness to visual distribution shift, and real-world deployment stays within seen embodiments.
Separately, several engineering details deserve to be remembered for their transferability: MRPE solves "multi-rate temporal alignment" with position encoding alone — a problem every video-action model hits; logical shards + merged frame requests cut video-training I/O by 30%; and the unified action space's handling of gripper direction, wheel velocities vs Cartesian commands is a standard pitfall checklist for cross-embodiment training.
7. Resources
- Paper: arXiv:2609.25627
- Project page: https://machembodied.com/ME-U/ME-U0.html
- Code: https://github.com/MachEmbodied/ME-U0