PAPER DEEP DIVE
One Policy, Many Embodiments: Unified Camera-Centric Action Geometry Pre-training for Heterogeneous Embodied Manipulation
UCAG-P from Xiaomi Embodied Intelligence and University of Macau unifies robot arms, bimanual platforms, humanoids and human-hand videos into one camera-centric action schema: 3D camera-frame anchor trajectories of the wrist (p0) and grasp center (p1), with a geometry-conditioned translator emitting executable commands — no retargeting or human-to-robot video synthesis needed. Built on Qwen3-VL-4B with three-stage training over 1.02M episodes / 6,373.6 hours / nine embodiments (36.7% human data), a single checkpoint without benchmark fine-tuning reaches 98.3% on LIBERO, 88.7%/89.2% on RoboTwin Easy/Hard, 82.0% zero-shot on LIBERO-Plus and 62.0% on RoboCasa GR-1; on real Piper robots bread pickup hits 60% vs 20% for pi0.5, with 35% zero-shot ALOHA-to-ARX transfer.
TL;DR
Xiaomi's Embodied Intelligence Team, together with University of Macau, introduces UCAG-P (Unified Camera-centric Action Geometry Pre-training): instead of sharing robot-specific joint commands across datasets, all manipulation behaviour is represented as camera-observable anchor motion — 3D camera-frame trajectories of the wrist/end-effector point $p_0$ and the grasp center $p_1$. Robot arms, bimanual platforms, humanoids, and even human-hand videos are all treated as different embodiments of one action schema, and a geometry-conditioned translator maps the shared geometric prediction into executable commands. A single checkpoint, with no benchmark-specific fine-tuning, reaches 98.3% on LIBERO, 88.7%/89.2% on RoboTwin Easy/Hard, 82.0% zero-shot on LIBERO-Plus, and 62.0% on RoboCasa GR-1.
1. The Problem: the Heterogeneity Wall
Scaling VLA policies is bottlenecked less by data volume than by data heterogeneity: the same behaviour appears as end-effector deltas, joint commands, gripper states, or human-hand keypoint sequences depending on the dataset. Existing remedies — explicit action retargeting, human-to-robot video synthesis, dataset-specific adaptation branches — fundamentally hinder joint training of a unified policy. Prior cross-embodiment systems (π0, GR00T, X-VLA, ABot-M0) still keep interfaces tied to robot-native commands or embodiment-specific heads, so human videos cannot supervise the policy as first-class citizens.
UCAG-P's key observation: the most stable common structure across embodied demonstrations lies not in the low-level controller, but in the camera-observable geometry of manipulation. Task progress is visible through the motion of the wrist and grasp center in the camera view — tied to the visual evidence a VLA uses, yet independent of any joint layout or controller.
2. Pre-training Data: 1,020,672 Episodes, 6,373.6 Hours, Nine Embodiments
The corpus spans eleven dataset subsets: 266.3 h of real-robot data (4.18%; DROID, RoboChallenge, RoboCoin-Piper), 3,767.5 h of simulation (59.11%, of which InternData-A1 synthetic data contributes 3,649.4 h), and 2,339.7 h of human-hand data (36.71%; EgoDex 829 h, EgoVerse 1,362 h, VITRA 148.7 h). A separate vision-language mixture (ShareRobot, RefSpatial-v2, RoboVQA, RoboAfford) adds instruction following, spatial grounding, affordance understanding, and trajectory reasoning. Human demonstrations are converted via MediaPipe keypoints: the wrist keypoint maps to $p_0$ and the thumb-index midpoint to $p_1$, yielding camera-centric pseudo-actions — no human-to-robot video synthesis or explicit retargeting required.
A four-stage pipeline unifies everything: episode indexing and domain partitioning (preserving dataset identity and mixture weights), visual/language standardization (224×224, fixed view slots, zero-filled missing views), geometric and action alignment ($p_0/p_1$ recovered via forward kinematics, calibration, and hand keypoints; robot-only quantities left invalid rather than synthesized), and temporal windowing (horizon $H=30$ with validity masks).
3. Method: Anchor Pairs, a Decoupled Architecture, Three-Stage Training
3.1 Camera-Centric Action Space
Each step is a 30-dimensional vector:
$$g_{t+h} = \left[\, m^L_{t+h},\; m^R_{t+h},\; \xi^{\text{cam}}_{t+h} \,\right] \in \mathbb{R}^{30}$$
Each manipulator $a \in \{L,R\}$ occupies a 10-dimensional slot:
$$m^a_{t+h} = \left[\, \Delta p^{\text{cam}}_{0,t+h},\; \Delta p^{\text{cam}}_{1,t+h},\; \Delta\psi^{\text{cam}}_{t+h},\; \gamma^{\text{gripper}}_{t+h},\; 0 \,\right] \in \mathbb{R}^{10}$$
where the two anchor displacements are measured relative to the chunk's first frame, $\Delta\psi$ is in-plane rotation (continuous sin/cos), and $\gamma$ is gripper state. The camera-motion term $\xi^{\text{cam}} = [\Delta t^{\text{cam}},\, r_6(\Delta R^{\text{cam}}),\, 0]$ activates only for moving cameras; single-arm samples mask the unused arm block; unavailable channels are masked out of the loss — each sample contributes only the targets it owns.
3.2 Architecture: Shared Prediction + Geometric Translation
The base policy $\pi_\theta(o_t, l) = \hat{g}_{t:t+H}$ uses a Qwen3-VL-4B-Instruct backbone; learnable action-query tokens feed a motion head (2560→5120 hidden + two residual MLP blocks) that predicts a 30-step, 30-dim shared geometric chunk. The translator conditions on:
$$\hat{u}^{(k)}_{t:t+H} = \phi_\psi\!\left(\, \hat{g}_{t:t+H},\; s^{(k)}_t,\; T^{(k)}_{\text{base}\leftarrow\text{cam}},\; J^{(k)}_t,\; z^{(k)}_{\text{geo}} \,\right)$$
with the camera-to-base transform, local Jacobian, and embodiment structure tokens, emitting an 80-dimensional sparse command layout — eight ordered 10-dim blocks (left/right arm joints, left/right end-effector XYZ+6D rotation, left/right hands, waist, mobile base); inactive blocks are masked and never dispatched. Eight learnable query tokens attention-pool VLM hidden states into the action head; this pooling lifts RoboCasa GR-1 from 58.3% to 62.0% (+3.7pp) in ablations.
3.3 Three-Stage Training
Stage 1 (camera-centric specialization): 128× H20, 200K steps, on all samples with camera-centric labels plus VLM data. Stage 2 (geometry-conditioned translation): 8× H20, 10K steps, on ground-truth trajectories to isolate the geometry→command mapping from upstream prediction error. Stage 3 (joint robot-human training): 64× H20, 10K steps, feeding the translator the policy's own predictions to match inference-time inputs, with generic VLM data excluded. Total objective:
$$\mathcal{L} = \mathcal{L}_{\text{geo}} + \lambda_{\text{cmd}}\, \mathcal{L}_{\text{cmd}}$$
both masked, normalized L1 losses, e.g.:
$$\mathcal{L}_{\text{geo}} = \frac{\sum_{h=1}^{H}\sum_d m^g_{h,d}\,\lvert \hat{g}_{t+h,d} - g_{t+h,d} \rvert}{\max\!\left(1,\; \sum_{h=1}^{H}\sum_d m^g_{h,d}\right)}$$

4. Results
| Benchmark | Setting | Best Specialist | UCAG-P (unified checkpoint) |
|---|---|---|---|
| LIBERO | Single-arm | Being-H0.7 99.2 / ABot-M0 98.6 | 98.3 |
| RoboTwin Easy | Bimanual | Being-H0.7 90.2 | 88.7 (+2.6 vs Qwen-VLA-Instruct) |
| RoboTwin Hard | Bimanual | Being-H0.7 89.6 | 89.2 (best Randomized average) |
| RoboCasa GR-1 | Humanoid | JoyAI-RA 63.2 | 62.0 (first on 6 of 24 tasks) |
| LIBERO-Plus | Zero-shot robustness | ABot-M0 80.5 | 82.0 (robot-state 92.8, lighting 98.9) |
Across the four standard LIBERO suites UCAG-P averages 98.25% with the best Goal score (99.2%) and only a 2.8-point spread, indicating consistent spatial reasoning, object interaction, goal following, and long-horizon execution. On LIBERO-Plus it is fully zero-shot — no LIBERO data used for post-training — yet still beats the specialist ABot-M0, with particular robustness to robot-state, lighting, and background perturbations; camera and sensor-noise shifts remain harder.
Real world: on Piper robots, bread pickup (evaluating human-to-robot transfer) reaches 60% vs π0.5's 20%, drawer opening 90% vs 85%, and bimanual bowl stacking 75% vs 65% under identical demonstration budgets, observations, and controllers. Cross-embodiment: replacing ALOHA with ARX in RoboTwin while holding scenes, perturbation seeds, and extrinsics fixed still yields 35.0% zero-shot — evidence that camera-centric motion is a transferable intermediate target, though morphology mismatch remains open.
In the 50-task RoboTwin breakdown, UCAG-P leads the Randomized average at 89.20% vs ZR-0's 87.98%; the clearest weakness is articulated opening (OpenMicrowave 11%/13%).

5. Method Pipeline at a Glance
1,020,672 episodes / 6,373 h
real + sim + human hand"] --> B["Unified data pipeline
indexing / standardization
geometric alignment / windowing"] B --> C["Stage 1: camera-centric specialization
Qwen3-VL-4B + motion head
200K steps / 128x H20"] C --> D["Stage 2: geometry-conditioned translator
ground-truth trajectories / 10K steps"] D --> E["Stage 3: joint robot-human training
translator consumes policy predictions / 10K steps"] E --> F["Single checkpoint"] F --> G["Shared output
30 steps x 30-dim anchor motion"] F --> H["Translator output
30 steps x 80-dim sparse commands"] G --> I["LIBERO 98.3 / RoboTwin 89.2
LIBERO-Plus zero-shot 82.0"] H --> J["Real Piper robots
bread pickup 60% vs pi0.5 20%"]
5. Limitations and Outlook
The authors acknowledge: (i) geometric supervision depends on reliable calibration and estimation — camera calibration, depth, kinematics, and MediaPipe keypoints errors propagate into targets and the translator; (ii) the ALOHA→ARX gap shows camera-centric motion alone does not eliminate morphological mismatch; (iii) real-world evaluation is deliberately controlled and limited in scale. Strategically, UCAG-P offers a route complementary to works like GE-Act 2.0: once the shared action space rises from robot commands to camera-observable geometry, human videos — the cheapest, most abundant manipulation data source — can directly supervise policy learning without expensive synthesis or retargeting. With 36.7% human data and a 3× real-robot bread-pickup gain over π0.5, the data dividend is already materializing. Project page and code: public-bots.github.io/UCAG-P.



