PAPER DEEP DIVE
Scaling Behavior Foundation Model for Humanoid Robots
Humanoid control requires natural whole-body coordination, precise real-time responses to control signals, and robust generalization across diverse environmental contexts, making it a cornerstone for generalist embodied agents. Behavior Foundation Models (BFMs) have recently emerged as a promising solution to address these challenges by leveraging large-scale behavioral data to achieve superior expressiveness, versatility and generalization. However, despite growing interest in scaling BFMs to further improve their capabilities, it remains unclear how key factors, including the learning paradigm, behavioral data and model architecture should be coordinated to enable effective scaling. In this work, we revisit the scaling recipe for BFMs and demonstrate that substantial performance gains can be achieved through the coordination of three core components: 1) the learning paradigm of motion tracking that reformulates diverse humanoid control problems as the reproduction of integrated whole-body behaviors in the global frame; 2) the strategic synergy between on-policy rollout quantity and reference motion diversity; and 3) the expressive and scalable model architecture termed Humanoid Transformer that facilitates the natural emergence of structured behavioral representations. Through extensive experiments in both simulation and real-world deployment, we demonstrate that our approach yields significant improvements in control fidelity and task generalization, reducing Mean Per-Keypoint Position Error (MPKPE) on the test set by over 10% in local mode and 82% in global mode compared with existing humanoid controllers. These results establish BFM as a principled and effective foundation for scalable and general-purpose humanoid control.
One-line summary
ScaleBFM systematically studies the scaling laws of Behavior Foundation Models for humanoid robots, revealing that three factors must be coordinated: motion tracking as unified learning paradigm, synergy between on-policy rollout quantity and reference-motion diversity, and the expressive scalable Humanoid Transformer architecture — reducing MPKPE by 10%+ in local mode and 82% in global mode on whole-body control.
Abstract
Humanoid control requires natural whole-body coordination, precise real-time responses to control signals, and robust generalization across diverse environmental contexts, making it a cornerstone for generalist embodied agents. Behavior Foundation Models (BFMs) address these by leveraging large-scale behavioral data. However, how to scale BFMs remains unclear — how should the learning paradigm, behavioral data, and model architecture be coordinated? This work revisits the scaling recipe and demonstrates substantial gains through three coordinated components: 1) motion tracking reformulating diverse humanoid control as reproduction of integrated whole-body behaviors in the global frame; 2) strategic synergy between on-policy rollout quantity and reference motion diversity; 3) the Humanoid Transformer architecture facilitating natural emergence of structured behavioral representations. Experiments in simulation and real-world deployment show significant improvements in control fidelity and generalization, reducing MPKPE by over 10% in local mode and 82% in global mode versus existing controllers.
1. Background and Motivation
Humanoid control is a cornerstone of general embodied intelligence but demands natural whole-body coordination (WBC), precise real-time responses to control signals, and robust cross-environment generalization. Behavior Foundation Models (BFMs) achieve expressiveness, versatility, and generalization through large-scale behavioral-data pretraining. Yet how to effectively scale BFMs lacks systematic study: how should the learning paradigm, behavioral data, and model architecture be coordinated?
Existing WBC systems tailor goal states and reward functions per task, failing to establish a unified abstraction across scenarios. ScaleBFM's starting point: use motion tracking as a unified, scalable proxy task that reformulates diverse behavior learning as imitation reproduction of reference motions, providing dense whole-body supervision and a measurable optimization objective. The key is not simply piling on data or parameters, but understanding each scaling factor's mechanism and coordinating them.
Figure 1: ScaleBFM enables a broad spectrum of diverse humanoid behaviors (agile locomotion, dexterous manipulation, coordinated loco-manipulation) in both simulation and the real world.
2. Core Method
2.1 Problem Formulation
Humanoid control is modeled as an MDP $M=\langle\mathcal{S},\mathcal{A},\mathcal{T},\mathcal{R},\gamma\rangle$. Each step's state $s_t\in\mathcal{S}$ comprises proprioceptive state $s_t^p$ and goal state $s_t^g$; action $a_t\in\mathcal{A}$ is desired joint angles executed by a low-level PD controller. Humanoid behavior is defined as a trajectory over proprioceptive states and actions (goal states are external specifications):
$$B\triangleq(s_0^p,a_0^p,s_1^p,a_1^p,\cdots,s_{T-1}^p,a_{T-1}^p,s_T^p).$$
BFMs decouple behavior learning from task-specific control modes, formulating pretraining as goal-conditioned RL maximizing expected return:
$$\max_\theta\mathbb{E}_{\tau\sim p_\theta(\tau)}\left[\sum_{t=0}^{T}\gamma^t r_t\right],\qquad p_\theta(\tau)=p(s_0)\prod_{t=0}p(s_{t+1}|s_t,a_t)\pi_\theta(a_t|s_t).$$
BFMs aim to: 1) identify a proxy task with a general reward shared across behaviors; 2) construct a versatile control interface from diverse behavioral specifications. Imitation is the natural paradigm, so whole-body motion tracking is the proxy task, with reward evaluating how faithfully the humanoid follows reference motions.
2.2 Learning Paradigm: Motion Tracking Recipe
Proprioceptive design: Asymmetric actor-critic in PPO. Actor proprioception $s_t^p\triangleq(\omega_t^{root},g_t,q_t,\dot{q}_t)$; critic may use privileged simulator info $\hat{s}_t^p\triangleq(h_t,p_t,\theta_t,v_t,\omega_t,q_t,\dot{q}_t)$.
Control interface: Masked whole-body target poses in root-relative Cartesian space, $x_t\triangleq(p_{t+1}^{ref}-p_t^{root},p_{t+1}^{ref}-p_t^{root},\theta_{t+1}^{ref}\ominus\theta_t^{root},\theta_{t+1}^{ref}\ominus\theta_t^{root})$. A link-wise mask sampled from $\mathcal{M}=\{m_0,\dots,m_7\}$ defines 8 control modes.
Reward design: The key distinction — does not remove root-position tracking or decouple root following from whole-body pose tracking, but requires reproducing reference motions as integrated whole-body trajectories in the global frame. Removing root tracking assigns nearly identical guidance to globally-distinct behaviors; decoupling root and pose tracking compromises coordination. Global-frame tracking provides coherent whole-body guidance with reduced ambiguity.
flowchart TB
RM["Reference Motions
102M frames @50FPS
LAFAN+AMASS+OMOMO+GRAB+
SnapMoGen+FineDance+BONES-SEED+Embody3D"] --> RT["Two-stage Retargeting
skeleton align + IK"]
RT --> REF["Humanoid reference motions"]
REF --> PPO["PPO Motion Tracking
(proxy task)"]
ENV["IsaacLab env
Unitree G1, 29-DoF"] --> PPO
PPO --> ROLLOUT["On-policy rollouts
(width=#GPUs, depth=horizon)"]
ROLLOUT --> ADAPT["Adaptive Sampling
up-weight failed motions"]
ADAPT --> REF
PPO --> BFM["Behavior Foundation Model
Humanoid Transformer"]
CTRL["Control modes x8
(masked whole-body targets)"] --> PPO
style PPO fill:#e0e7ff,stroke:#2563eb
style BFM fill:#dcfce7,stroke:#16a34a
style ADAPT fill:#fef3c7,stroke:#d97706
2.3 Training Data: Quantity-Diversity Synergy
Key insight: under PPO, effective training data is on-policy rollouts whose scale depends on environment parallelism and rollout horizon; increasing reference motions shapes the rollout distribution, and its benefit depends critically on the diversity of the reference corpus rather than size alone. An arbitrarily large set of forward-walking motions represents only a single behavior pattern.
Figure 3: Reference motion dataset composition. Aggregates 102M frames @50FPS from multiple sources, retargeted to the target humanoid.
Adaptive sampling: Each sequence starts with $w_0=1$; every $T_{eval}$ epochs the model is evaluated on the full training set. A trajectory's weight is updated:
$$w_t=\begin{cases}w_{t-1}/\beta^{T_{eval}},&\text{if trajectory fails}\\w_{t-1}\cdot\beta^{T_{eval}},&\text{if trajectory succeeds}\end{cases},\qquad w_t\leftarrow\text{clip}(w_t,w_{min},w_{max}).$$
In practice $T_{eval}=200,\beta=0.999,w_{min}=0.03,w_{max}=1.0$, forming a closed loop of evaluation-feedback-resampling focusing training on difficult motions while retaining coverage.
2.4 Humanoid Transformer Architecture
Proprioception, goals, and actions are extended into finite temporal windows, encoded by modality-specific tokenizers. Proprioception and action tokens are interleaved into the context sequence, followed by a learnable query token. In self-attention the query token is not attended to by context tokens but retains full attention over the context to aggregate history. Goal states with temporal offsets are concatenated and injected via cross-attention.
Figure 2: Humanoid Transformer architecture. RMSNorm normalizes goal embeddings onto a continuous bounded hypersphere, naturally inducing a structured latent representation space for behavioral intentions.
Structured latent space: RMSNorm normalizes goal embeddings onto a hypersphere without explicit regularization — shaped solely by the motion-tracking objective. The actor's future window has 5 consecutive frames $\{0,1,2,3,4\}$ plus one random frame from $[5,32]$ (mitigating deployment delay); the critic uses exponentially-spaced $\{0,1,2,4,8,16,32\}$ for both short-term detail and long-horizon outcomes.
3. Key Experiments
Trained on Unitree G1 (1.3m, 29 DoF) in IsaacLab, evaluated in MuJoCo for dynamic-transfer robustness. Two test sets: BONES (10k held-out sequences, ~3.6M frames) and Ours (839 Xsens + 810 100Style sequences, ~4.5M frames).
3.1 Scaling On-Policy Data
Jointly scaling width (GPU count) and depth (horizon) substantially improves performance; the largest config (64 GPUs + horizon 64) is best in nearly all experiments. Scaling either dimension alone does not consistently improve — the benefit depends not only on total experience but on how it's accumulated, requiring appropriate width-depth balance.
3.2 Scaling Reference Motions: Homogeneous vs Heterogeneous
Behavioral coverage is quantified via occupancy rate. XXS→S keeps occupancy ~constant (~0.936); S→L rises sharply (0.937→0.9995). This splits reference-motion scaling into two regimes: homogeneous (XXS→S, quantity up but coverage flat) yields only marginal gains; heterogeneous (S→L, quantity and coverage both up) brings little benefit on BONES but large gains on Ours. Effectiveness depends on whether scaling expands coverage and whether new behaviors are relevant to the target benchmark.
| Partition | Occupancy Rate | Coverage |
|---|---|---|
| XXS | 0.9365 | ~flat |
| XS | 0.9350 | ~flat |
| S | 0.9365 | ~flat |
| M | 0.9825 | improved |
| L | 0.9995 | greatly improved |
Table 1: Training partition occupancy rates. XXS→S homogeneous, S→L heterogeneous.
3.3 Scaling Model Architecture
The Humanoid Transformer generally achieves higher success rates and lower tracking errors than the MLP. A medium Transformer already matches or exceeds a much larger MLP; further scaling yields diminishing returns. Scaling is inconsistent across control modes — whole-body is easier to learn than root-and-end-effector (which infers whole-body behavior from sparse constraints), and optimization trade-offs between modes intensify as training progresses.
3.4 Latent Space Analysis
The latent space satisfies locality (temporally close commands map nearby) and global organization (different-direction behaviors occupy distinguishable regions). Visualization shows forward/backward walking on opposite ends of one dimension, crouch-left/right separated along another. Moderate latent noise is tolerated with strong performance, proving robustness.
Figure 4: Latent space visualization. Each behavior forms a continuous smooth trajectory; different-direction behaviors occupy distinguishable regions preserving directional relationships.
3.5 Performance Benchmark
Under whole-body control, the 3M-parameter ScaleBFM leads GMT, TWIST, and SONIC in success rate and tracking errors on both test sets. Versus BFM-Bym (BeyondMimic-reward ablation), global tracking error is lower, proving the global-frame integrated-tracking reward provides more coherent behavioral guidance.
| Test Set | Method | Success Rate↑ | G-MPKPE↓ | L-MPKPE↓ |
|---|---|---|---|---|
| BONES | SONIC | 0.9239 | 0.1740 | 0.0436 |
| BFM-Bym (ablation) | 0.9644 | 0.1005 | 0.0406 | |
| BFM-Global | 0.9677 | 0.0798 | 0.0400 | |
| Ours | SONIC | 0.5937 | 0.5035 | 0.0430 |
| BFM-Bym (ablation) | 0.9709 | 0.1224 | 0.0403 | |
| BFM-Global | 0.9776 | 0.0915 | 0.0396 |
Table 2: Whole-body control benchmark. BFM-Global is best on both sets. G-MPKPE = global mean per-keypoint position error (lower is better).
4.
GCRL目标函数
$$ \max_{\theta}\,\mathbb{E}_{\tau\sim p_{\theta}(\tau)}\left[\sum_{t=0}^{T}\gamma^{t}r_{t}\right] $$
掩码目标状态
$$ s_{t}^{g}\triangleq(x_{t}\odot m_{i},\,m_{i}),\qquad m_{i}\sim\mathcal{M} $$
Limitations and Future WorkAuthor-stated limitations:
- Although 8 control modes are designed, whether this is the most appropriate abstraction and how these modes integrate with future high-level policies remains unclear; BFM design should evolve with high-level policy advances.
- The current BFM pretraining infrastructure remains preliminary and constrained, limiting scalability to broader settings.
Analysis: ScaleBFM's core contribution is not a new method but systematically clarifying "what to scale and why it works" — especially the homogeneous vs heterogeneous distinction (more quantity ≠ more coverage). This directly guides dataset curation: in-domain gains need fine-grained coverage not more volume; general capability gains need complementary data sources. The structured latent space's natural emergence and convergence with scale suggest larger models may unify representations across control modes. But the 29-DoF Unitree G1 is a small skeleton; transfer to higher-DoF or different-morphology humanoids is unverified; on-policy data scaling needs many GPUs (64-GPU config) posing reproduction barriers; inter-mode optimization trade-offs remain open.
5. Conclusion
The core idea of ScaleBFM is that scaling BFMs is not about piling on data and parameters, but coordinating the learning paradigm, behavioral data, and model architecture — especially distinguishing "quantity scaling" from "coverage scaling". It does three things: uses global-frame motion tracking as a unified proxy task for coherent guidance; balances quantity and diversity through on-policy width-depth synergy plus adaptive sampling; and lets structured behavioral representations emerge naturally via the Humanoid Transformer's hyperspherical latent space. Empirically, the 3M-parameter model leads on whole-body control across both test sets, with global MPKPE reduced by 10%+ (local) and 82% (global). The work provides a foundation for systematic BFM development, though control-interface appropriateness, infrastructure scalability, and cross-morphology transfer remain open.
SOURCE LINKS



