PAPER DEEP DIVE
CMoE: Contrastive Mixture of Experts for Motion Control and Terrain Adaptation of Humanoid Robots
CMoE uses contrastive learning to improve MoE expert specialization for humanoid terrain navigation. Validated on Unitree G1, achieving robust gait over 20cm steps and 80cm gaps.
One-sentence summary: CMoE imposes a contrastive learning constraint on the gate outputs of a mixture-of-experts policy, allowing a single Unitree G1 policy to traverse 20 cm continuous steps, 80 cm gaps, 30 cm hurdles, and 17° slopes while avoiding the near-uniform expert activation that limits vanilla MoE.
Paper information: CMoE: Contrastive Mixture of Experts for Motion Control and Terrain Adaptation of Humanoid Robots, accepted at ICRA 2026. The work is by Shihao Ma, Hongjin Chen, Zijun Xu, Yi Zhao, Ke Wu, Ruichen Yang, Leyao Zou, Zhongxue Gan, and Wenchao Ding from Fudan University.
Links: arXiv · Project page · Code repository · Video
Background and Motivation
Humanoid deployment in real environments requires controllers that handle continuous terrain sequences with abrupt transitions. Flat ground can turn into stairs, gaps, hurdles, slopes, and composite courses within a few meters. A successful policy must maintain balance while also changing foot placement, body posture, and swing-leg behavior at short spatial scales. Compared with quadrupedal robots, humanoids have a higher center of mass and a narrower support polygon, so terrain misclassification can quickly lead to falls.
Recent learning-based systems often combine terrain perception and motion control in a single network. Mixture-of-experts is a natural architecture for this problem because it can assign specialized subnetworks to different skills while a gate selects or blends expert outputs. MoE has been used widely in autonomous driving, natural language processing, computer vision, and legged locomotion, and it is often viewed as a way to reduce gradient conflicts and task interference in multi-task RL.
The paper identifies a practical failure of vanilla MoE: lazy gating. Although the architecture is expressive enough to model multiple terrain regimes, the gating network tends to activate all experts with nearly equal weights regardless of terrain. The MoE then degenerates into a large monolithic network, and expert specialization does not emerge. The authors argue that expert division of labor should be learned explicitly rather than assumed by the architecture.
CMoE addresses this problem by aligning gate outputs with terrain features in a shared contrastive learning space. Pairs of gate and terrain encodings from the same trajectory are treated as positives, while other pairs are treated as negatives. A SwAV-style objective with Sinkhorn-Knopp cluster assignments prevents trivial solutions and drives the gate to produce terrain-specific activation patterns.
Method
System Overview
CMoE is a single-stage reinforcement learning framework with three interacting components: context-state distillation, a mixture-of-experts policy, and a terrain contrastive objective. The distillation module extracts an explicit body-state estimate and implicit latent representations from proprioceptive history and an elevation map. The MoE actor-critic then generates actions and value estimates from those representations. The contrastive loss ties the gate outputs to the encoded terrain semantics.
Problem Definition and PPO
The authors formulate multi-terrain locomotion as a Markov decision process with a state space, action space, transition distribution, reward function, and discount factor. The policy is trained with proximal policy optimization by maximizing the expected discounted return:
$$\pi^{*}=\arg\max\mathbb{E}\left[\sum_{t=0}^{\infty}\gamma^{t}R(s_{t},a_{t})\right]$$
The Unitree G1 policy outputs 12-dimensional leg actions, which are converted into joint targets by a proportional controller. Training is performed in Isaac Gym with 4096 parallel environments. Unlike two-stage teacher-student pipelines, CMoE optimizes perception, gating, experts, and value functions in one RL update loop.
Observation Design
The proprioceptive observation contains angular velocity, gravity direction, velocity command, joint angles, joint velocities, and the previous action:
$$\mathbf{o}_{t}=[\omega_{t},g_{t},c_{v}^{t},\theta_{t},\dot{\theta}_{t},a_{t-1}]$$
In the implementation, each one-step observation is 45-dimensional, and the network uses the last 10 steps to form a 450-dimensional proprioceptive history. The elevation map contains 77 sampled height points covering roughly a 0.7 m × 1.1 m rectangle in front of the robot. This design gives the policy both short-term motion history and local terrain geometry.
Context-State Distillation
The context-state distillation model contains two estimators. The first is a β-VAE that infers base linear velocity and a compact latent state from proprioceptive history. The second is an autoencoder that compresses the elevation map into low-dimensional terrain features. Both estimators are trained with self-supervised reconstruction or prediction objectives and do not require external labels.
The state estimator loss combines velocity prediction with the VAE objective:
$$\mathcal{L}_{\mathrm{CS}}=\mathrm{MSE}(\tilde{\mathbf{v}}_{t},\mathbf{v}_{t})+\mathcal{L}_{\mathrm{VAE}}$$
The velocity error uses simulator ground truth. The VAE component contains next-step observation reconstruction and a KL regularizer on the latent distribution:
$$\mathcal{L}_{\mathrm{VAE}}=\mathrm{MSE}(\tilde{\mathbf{o}}_{t+1},\mathbf{o}_{t+1})+\beta D_{\mathrm{KL}}(q(\mathbf{z}^{H}_{t}\mid\mathbf{o}^{H}_{t})\parallel p(\mathbf{z}^{H}_{t}))$$
The terrain estimator uses a lighter autoencoder with a reconstruction loss:
$$\mathcal{L}_{\mathrm{AE}}=\mathrm{MSE}(\tilde{\mathbf{e}}_{t},\mathbf{e}_{t})$$
The code for these estimators is located at `rsl_rl/rsl_rl/modules/state_estimator.py` and `rsl_rl/rsl_rl/modules/terrain_estimator.py`. The state estimator implements reparameterized VAE sampling, while the terrain estimator uses the encoder mean as the terrain latent. Their updates are performed before the PPO policy update, so gating and experts always consume fresh explicit and implicit context.
Mixture-of-Experts Policy
CMoE applies the MoE structure to both the actor and critic. Each expert is an independent actor-critic pair, and the same gating network is shared between actor and critic so that action generation and value estimation use consistent expert weights. The final action mean is a softmax-weighted combination of expert outputs:
$$\mu_{\mathrm{weighted}}=\sum_{i=1}^{N}\mathrm{softmax}(g_{i})\cdot\mu_{i}$$
The gate input is built by concatenating the current observation, explicit body state, proprioceptive latent, and terrain latent. In `cmoe_actor_critic.py`, the gate is a two-layer MLP followed by softmax, and the five experts share the same fused input but maintain separate parameters. In the critic evaluation path, expert values are weighted by detached gate weights to avoid the value branch updating the gate indirectly.
# rsl_rl/rsl_rl/modules/cmoe_actor_critic.py
expert_meansm = torch.stack(expert_means, dim=1)
weighted_means = (expert_meansm * self.gate_weights.unsqueeze(2)).sum(dim=1)
self.distribution = Normal(weighted_means, self.std)
return self.distribution.sample()
Shared gating keeps action and value semantics aligned. The experiments later show that, without contrastive learning, gate weights barely change with terrain. This suggests that the PPO reward alone is not enough to force stable expert differentiation.
Terrain Contrastive Learning
The contrastive objective maps gate outputs and terrain encodings into the same 16-dimensional space through two projection heads, producing $g^{z}_{t}$ and $e^{z}_{t}$. Following SwAV, the model computes cluster assignment probabilities over L2-normalized learnable prototypes:
$$\mathbf{p}^{g}_{t}=\frac{\exp(\frac{1}{\tau}{g^{z}_{t}}^{\top}e_{k})}{\sum_{k'}\exp(\frac{1}{\tau}{g^{z}_{t}}^{\top}e_{k'})}$$
To avoid collapse to a single cluster, Sinkhorn-Knopp produces target assignments $q$. The contrastive loss then minimizes the cross prediction error in both directions:
$$\mathcal{J}^{\mathrm{SwAV}}=-\frac{1}{2H}\sum^{H}_{t=1}(q^{g}_{t}\log p^{e}_{t}+q^{e}_{t}\log p^{g}_{t})$$
The implementation uses 32 prototypes, a temperature of 0.2, three Sinkhorn iterations, and epsilon 0.05. Both projection heads use 128-64-16 MLPs. The contrastive loss is added to the PPO loss and backpropagated together with the policy surrogate loss, while the gate input and height input are detached to prevent interference with the estimator internals.
# rsl_rl/rsl_rl/modules/cmoe_actor_critic.py
score_s = gate_z @ self.prototypes.weight.T
score_t = height_z @ self.prototypes.weight.T
with torch.no_grad():
q_s = sinkhorn(score_s)
q_t = sinkhorn(score_t)
log_p_s = F.log_softmax(score_s / self.temperature, dim=-1)
log_p_t = F.log_softmax(score_t / self.temperature, dim=-1)
contrastive_loss = -0.5 * (q_s * log_p_t + q_t * log_p_s).mean()
The following Mermaid diagram summarizes the data flow from state encoding, through gating and expert blending, into the contrastive and PPO losses:
flowchart TD
A[Proprioceptive history 450 dims] --> B[StateEstimator beta-VAE]
B --> C[Explicit body state]
B --> D[Latent state z_H]
E[Elevation map 77 points] --> F[TerrainEstimator AE]
F --> G[Terrain latent z_E]
H[Current observation] --> I[Fused actor input]
C --> I
D --> I
G --> I
I --> J[Gating network]
I --> K[Five ExpertActorCritic]
J --> L[Gate weights]
K --> M[Weighted action mean]
L --> M
L --> N[GateProjector]
G --> O[TerrainProjector]
N --> P[SwAV prototype clustering]
O --> P
P --> Q[Contrastive loss J_SwAV]
M --> R[IsaacGym / Unitree G1]
Q --> S[PPO total loss]
S --> R
Domain Randomization and Elevation Noise
To bridge the sim-to-real gap, the authors randomize joint mass, moment of inertia, friction, restitution, motor strength, and motor gains $k_p$ and $k_d$. The robot receives external perturbations of up to 30 N every 16 seconds. The elevation map is corrupted with delay noise, Gaussian noise, and randomized offset and rotation.
The paper introduces nonlinear salt-and-pepper noise for unstable height extremes. Each height point $h(i)$ is replaced with a uniform sample near the local maximum or minimum with probability $p$, and retained otherwise:
$$h(i)=\begin{cases}\mathcal{U}(M,2M-m) & \text{with probability }p,\\ \mathcal{U}(2m-M,m) & \text{with probability }p,\\ h(i) & \text{otherwise}\end{cases}$$
Here $M$ and $m$ are the local maximum and minimum values around the corresponding height point. The paper also chamfers sharp right-angle edges in simulated elevation maps because real sensors often produce smooth curved edges. Together these operations improve robustness to noisy elevation perception during deployment.
Training and Implementation Details
Training uses a single NVIDIA RTX 4090, 4096 parallel environments, and 20,000 epochs. The policy uses five experts, 32 prototypes, a temperature of 0.2, and an elevation map covering 0.7 m × 1.1 m. Curriculum learning divides terrains into simple and complex groups and gradually increases velocity-command magnitude and direction on complex terrains.
The official repository defines the full environment configuration in `legged_gym/legged_gym/envs/g1/g1_cmoe_config.py`. The curriculum simultaneously trains on flat ground, rough slopes, stairs, discrete protrusions, gaps, hurdles, mixed terrains, and narrow stairs, with terrain sampling controlled by `terrain_proportions`:
# legged_gym/legged_gym/envs/g1/g1_cmoe_config.py
terrain_dict = {
"plane": 0.,
"rough slope": 0.1,
"stairs up": 0.1,
"stairs down": 0.1,
"discrete": 0.1,
"parkour_gap": 0.3,
"parkour_step_up": 0.,
"parkour_step_down": 0.,
"parkour_hurdle": 0.1,
"mix": 0.1,
"narrow_stairs": 0.1,
"composite": 0.,
}
terrain_proportions = list(terrain_dict.values())
The reward design is summarized in Table II of the paper. Positive rewards emphasize velocity and yaw tracking; negative terms penalize collisions, excessive base height, joint limit violations, high joint velocity or torque, rapid action changes, and foot stumbling. The foot-edge reward is enabled only on hurdle or gap terrain because contacting step edges should not always be penalized.
| Term | Expression | Weight |
|---|---|---|
| Velocity tracking | $\exp\{-\|\mathbf{v}_{xy}-\mathbf{v}_{xy}^{c}\|_{2}^{2}/\sigma\}$ | 2.0 |
| Yaw tracking | $R_{\mathrm{yaw}}=\exp(-|\psi_{\mathrm{cmd}}-\psi|)$ | 2.0 |
| Z velocity | $\mathbf{v}_{z}^{2}$ | -1.0 |
| Roll-pitch velocity | $\|\boldsymbol{\omega}_{xy}\|_{2}^{2}$ | -0.05 |
| Orientation | $\|\mathbf{g}_{x}\|_{2}^{2}+\|\mathbf{g}_{y}\|_{2}^{2}$ | -2.0 |
| Base height | $(h-h^{\mathrm{target}})^{2}$ | -15.0 |
| Feet stumble | $\bigvee_{i\in\text{feet}}\{\|\mathbf{F}_{i,xy}\|_{2}>3\cdot|F_{i,z}|\}$ | -1.0 |
| Collision | $\sum_{i\in\mathcal{I}_{\mathrm{penalty}}}\mathbb{I}(\|\mathbf{F}_{i}\|_{2}>0.1)$ | -15.0 |
| Action rate | $\|\mathbf{a}_{t}-\mathbf{a}_{t-1}\|_{2}^{2}$ | -0.3 |
| Joint velocity | $\sum_{i}\dot{\theta}_{i}^{2}$ | -5.0e-4 |
| Torque | $\sum_{i}\tau_{i}^{2}$ | -1.0e-5 |
| Joint position limits | $\mathrm{ReLU}(\theta-\theta_{\mathrm{max}})+\mathrm{ReLU}(\theta_{\mathrm{min}}-\theta)$ | -2.0 |
The training terrain ranges are shown below. The simulation benchmark uses a 3 m × 18 m runway, a command speed of 0.8 m/s, and a 20-second episode. A trial fails if any part other than the feet collides or if torso roll or pitch deviation exceeds 1°.
| Terrain | Description | Range |
|---|---|---|
| Slope | Slope angle | 0-20° |
| Stairs | Step height | 0.05-0.23 m |
| Gap | Ditch width | 0.1-0.8 m |
| Hurdle | Height / width | 0.2-0.4 m / 0.1-0.3 m |
| Discrete | Irregular protrusion height | 0.1-0.2 m |
| Mix1 | Gaps and steps, gap width / step height | 0.1-0.8 m / 0.1-0.15 m |
| Mix2 | Single-log bridge and steps, bridge width / step height | 0.5-1.0 m / 0.1-0.25 m |
Experimental Analysis
Simulation Performance
The authors compare CMoE with two baselines. Base is an actor-critic without MoE but with the same parameter count. Vanilla MoE uses a basic MoE structure without the terrain encoder and contrastive learning. The metrics are success rate over the full course and average travel distance along the forward direction within the time limit.
| Method | Slope | Stairs Up | Stairs Down | Discrete | Gap | Hurdle | Mix1 | Mix2 |
|---|---|---|---|---|---|---|---|---|
| CMoE | 0.991 | 0.886 | 0.905 | 0.991 | 0.974 | 0.987 | 0.767 | 0.747 |
| Vanilla MoE | 0.957 | 0.798 | 0.908 | 0.987 | 0.818 | 0.970 | 0.605 | 0.662 |
| Base | 0.966 | 0.481 | 0.483 | 1.000 | 0.221 | 0.779 | 0.276 | 0.388 |
| Method | Slope | Stairs Up | Stairs Down | Discrete | Gap | Hurdle | Mix1 | Mix2 |
|---|---|---|---|---|---|---|---|---|
| CMoE | 14.870 | 10.802 | 10.824 | 13.440 | 14.876 | 13.470 | 12.055 | 9.750 |
| Vanilla MoE | 11.675 | 8.898 | 9.250 | 14.870 | 11.980 | 14.780 | 9.960 | 8.703 |
| Base | 12.917 | 8.210 | 8.726 | 15.280 | 7.385 | 10.124 | 8.209 | 8.848 |
CMoE maintains high success rates on slopes, stairs, gaps, hurdles, and mixed terrains. Base reaches 1.000 on discrete terrain but only 0.221 on gaps, showing that a single dense network struggles to cover both discrete foothold selection and large gap crossing. Vanilla MoE approaches CMoE on several individual terrains but drops substantially on gaps and mixed courses, which aligns with the missing contrastive learning.
Average travel distance follows a similar pattern. CMoE reaches 14.876 m on gaps, 12.055 m on Mix1, and 9.750 m on Mix2. Vanilla MoE is slightly higher on hurdles at 14.780 m but falls to 9.960 m and 8.703 m on Mix1 and Mix2. Expert specialization appears most valuable when the terrain switches rapidly within one episode.
Effect of Contrastive Learning
The t-SNE analysis shows that Vanilla MoE activation points do not separate across terrains. In contrast, CMoE activation points form terrain-specific clusters, with similar terrains grouped and stratified by difficulty. A notable result is that ascending stairs and descending stairs are far apart in representation space. The model does not treat all stair tasks as one category; it learns the different dynamics of lifting the swing leg versus controlling descent.
Training curves show that CMoE reaches higher terrain levels earlier than Base and achieves faster reward growth with a higher peak. The authors attribute this to terrain classification performed by the gating network, which reduces multi-task interference and accelerates convergence.
Expert Behavior Analysis
The authors build a test course with three 30 cm hurdles, a 15° uphill slope, ten downhill steps, and three 60 cm gaps. They record expert activation over time. Vanilla MoE keeps each expert in a small fixed activation range with little variation when terrain changes. CMoE produces clear activation jumps at terrain boundaries, with different experts taking different roles.
The paper also removes Expert 1 from the evaluation policy. When crossing stairs, the robot attempts to lift its leg but fails and falls; in downhill scenarios it still walks normally. This provides behavioral evidence that Expert 1 specializes in ascent-style foot lifting rather than global balance or descent control. This is stronger validation than activation plots alone.
Real-World Experiments
CMoE was deployed on a Unitree G1. The real system uses radar to collect point clouds and fuses them with the robot localization system to generate elevation-map observations. The policy then receives the elevation map as input and runs directly in the real world without additional fine-tuning.
The robot crossed continuous steps up to 20 cm high, gaps up to 80 cm wide, hurdles up to 30 cm high, and slopes of 17°. In a mixed course, it climbed 15 cm steps, crossed a 60 cm gap, cleared a 30 cm hurdle, and climbed an uphill slope. Robustness tests included untrained outdoor step edges, rope dragging, and object collisions. When struck by an object while crossing a 30 cm obstacle, the robot maintained stability and completed the task.
These results suggest that CMoE transfers well because it couples two mechanisms: elevation-aware state estimation gives the policy a look-ahead terrain prior, and contrastive learning stabilizes gate responses around terrain semantics. When the terrain changes, the gate can switch expert usage without waiting for a failure signal from the reward.
Code and Reproducibility
The official repository is released under the BSD-3-Clause license and includes the Isaac Gym environment, G1 configuration, CMoE actor-critic, PPO algorithm, terrain generators, and training or playback scripts. The README provides installation and usage instructions. The basic training command is:
# README.md
python legged_gym/legged_gym/scripts/train.py --task=g1cmoe --alg=cmoe --run_name cmoe
The initial release contains simulation training code only. The README and TODO.md state that MuJoCo playback, Unitree SDK deployment, and pretrained checkpoints are planned for future releases. The project page and video provide qualitative visual evidence for the real-world results.
Limitations
- Limited quantitative real-world evidence: The 20 cm step, 80 cm gap, 30 cm hurdle, and 17° slope results are presented as successful demonstrations. The paper does not report repeated-trial success rates, travel distances, or statistical variability for real-robot experiments.
- Incomplete code release: The repository currently provides only Isaac Gym simulation training code. MuJoCo evaluation, real-robot deployment, and pretrained checkpoints remain on the roadmap, so the full training-to-deployment pipeline is not yet directly reproducible.
- Limited ablation coverage: The main comparison is CMoE versus Vanilla MoE versus Base. The paper does not systematically ablate prototype count, temperature, expert count, projection dimension, or the separate contributions of state distillation and contrastive learning to sim-to-real transfer.
- Single platform and robot: All experiments use Unitree G1. The method is not evaluated on humanoids with different sizes, joint configurations, or controllers, so the reported limits may be coupled to G1-specific dynamics.
- Strict termination condition: Simulation trials terminate when torso roll or pitch deviation exceeds 1°. This strict threshold can lower success rates, but it also encourages the policy to learn conservative posture recovery.
Conclusion and Outlook
CMoE contributes more than a new MoE architecture for humanoid control. It addresses the structural failure that prevents vanilla MoE from realizing its theoretical expressiveness: experts do not specialize. By projecting gate outputs and elevation encodings into a shared prototype space, the model learns terrain-dependent expert activation and provides behavioral evidence that experts perform different motor functions.
From a systems perspective, CMoE preserves the simplicity of single-stage RL. Context-state distillation supplies body velocity and terrain latents, the MoE decomposes the control task across action generation, and contrastive learning binds terrain semantics to gate selection. These components co-evolve in the PPO update loop, enabling high success rates on difficult individual terrains and continuous behavior switching on mixed courses.
The authors identify whole-body control as the next step, aiming for coordinated arm, torso, and leg behavior in parkour-style tasks. For the community, the most valuable future additions would be deployment code, pretrained checkpoints, multi-platform evaluation, and a more detailed quantitative study of real-world robustness.
Golden Quote
Expert specialization is not the default result of a mixture-of-experts architecture; it must be learned. When gate outputs and terrain encodings are aligned in the same clustering space, the robot learns to call the right skill for the terrain it sees.



