
Introducing PROWL-2: Fidelity-Gated Dual Curricula — Jointly Learning Simulation and Decision-Making
Odyssey (with UCL AI Centre and the University of Basel) released PROWL-2 on Oct 1: the first framework coupling an agent curriculum and a world-model repair curriculum inside one continual training loop. Core insight: in imagination training, a high-learning-signal trajectory is ambiguous — a real policy weakness or just a wrong world-model prediction; prioritizing it indiscriminately reinforces the model's own hallucinations. A fidelity gate separates 'useful' from 'trustworthy': reliable high-potential imagined experience feeds the policy curriculum, unreliable rollouts — with their already-stored real continuations — go to a repair pool, fixed by a randomly-initialized, KL-anchored developer policy exploring the real environment, re-audited and readmitted once repaired. First place on all nine SMACv2 scenarios (+4-18% at 5v5, +20-91% at 10v10/10v11 over the backbone); big margins on hard MQE tasks — Gate-3 70.4% vs 29.8%, Shepherd-Hard 28.2% vs 7.3%, all other baselines at 0. Ablations show the curricula are not additive: repair alone is marginal, the ungated curriculum falls below the backbone in all nine scenarios, and the gate turns the same curriculum into the largest single gain. Architecture- and algorithm-agnostic; next steps point at humanoid coordination and long-horizon multi-agent games.
What World Models Promise — and Where the Promise Breaks
Multi-agent reinforcement learning (MARL) carries an expensive legacy: coordination is hard to learn. Partial observability forces every agent to act on incomplete information, and inter-agent dependencies blow up the joint state–action space as the team grows. Model-free methods (the QMIX / MAPPO family) typically burn enormous amounts of real environment interaction before coordinated behavior emerges.
World models are the community's prescription: learn the environment's dynamics and let the policy train on imagined rollouts instead of real trial and error. Dreamer-style learners and a growing family of multi-agent world models — MAMBA, CoDreamer, MARIE, DIMA, SeqWM — have carried these sample-efficiency gains into cooperative MARL.
But once trajectories can be generated on demand, a new question surfaces immediately: which pieces of imagined experience are worth learning from?
The Core Ambiguity: High Learning Signal ≠ Good Experience
Curriculum methods (prioritized replay, PLR, UED) select experience by learning potential — typically a TD-error or value-loss proxy: where the policy predicts poorly is where it has something to learn. That premise holds in the real environment, where experience comes from the environment itself. Inside a world model it breaks. A high-error imagined trajectory is ambiguous:
- It may expose a genuine weakness of the task policy — exactly what you want the policy to practice.
- Or it may simply be a low-fidelity rollout — dynamics that would never occur in the environment. Training on it teaches the policy to treat the model's hallucinations as reality.
Compounding model error is a classic failure mode of learning in imagination. Prior work fixed one side each: IMAC prioritizes imagined experience by learning potential but assumes a frozen world model during policy learning; PROWL-1 actively discovers and repairs world-model failures with an adversarial developer policy but maintains no curriculum over what the policy trains on. The two obviously depend on each other — model errors matter most exactly where the agent trains, and training is reliable only where the model is accurate — yet nobody had coupled them.

Fidelity-gated dual curricula: imagined rollouts pass a fidelity gate first; trustworthy high-potential experience enters the agent curriculum, unreliable rollouts — together with their recorded continuations — go to the world-model repair pool, and return after repair and re-audit (diagram: this site, after the paper)
PROWL-2: One Training Loop, Two Curricula, One Gate
PROWL-2, published October 1 by Odyssey (Ahmet H. Güzel, Jenny Seidenschwarz, Jeffrey Hawke, with the UCL AI Centre and the University of Basel), is the first framework to put both problems into a single continual training loop. Three learners are involved:
- Task policy $\pi_G$ — the multi-agent policy we want to train; it learns entirely inside the world model's imagined rollouts.
- Developer policy $\pi_D$ — never acts for the task; explores its own instance of the real environment searching for states where the world model fails.
- World model $M_\theta$ — itself a learner: it keeps training on ordinary replay, and is additionally repaired on the failures the developer finds and those the task policy runs into.

PROWL-2 overview: task policy curriculum (top left), world-model curriculum (top right), coupled through the fidelity gate (bottom) (source: paper Figure 1)
Curriculum One: Where the Policy Practices
IMAC's prioritization, extended to multi-agent: a start state $\ell=(e,t)$ in the curriculum buffer $\mathcal{B}_G$ indexes the joint action–observation history of all agents at step $t$ of a recorded episode, and imagination rolls the whole team forward together. Learning potential reuses IMAC's positive value loss averaged over agents (for the SMACv2 actor–critic backbone); on MQE's planner backbone it is measured as the magnitude of the critic's TD residual $|\delta Q|$ on real replay. Batches are drawn by PLR rank-based prioritization with staleness, plus uniform exploration.
Curriculum Two: Where the Model Gets Repaired
The developer policy is trained with PPO on an objective (Eq. 2) that adds a KL anchor pulling $\pi_D$ toward its random initialization $\pi_{ref}$. The detail is counterintuitive: not copying the task policy at initialization yields a higher task win rate (confirmed by ablation), and the KL anchor prevents the entropy collapse seen without it. Unlike UED adversaries that hunt for the task policy's failures, the developer hunts for the world model's failures.
The failure score $s_{fail}$ (Eq. 4) deliberately avoids a single reconstruction loss in favor of control-relevant errors: on SMACv2, ally death timing, available-action-mask disagreement, reward error and observation error; on MQE, per-agent latent errors at horizons 1–3 plus reward error. Each is standardized against running statistics and averaged with equal weights. The reward (Eq. 3) includes a learning-progress term $\beta_{LP}(s^{before}-s^{after})$ — if an error does not shrink after repair, the search steers away from it, because it is an error the model cannot learn to fix.
The Gate: What Can Be Trusted
This is PROWL-2's genuinely novel component, with an elegant engineering trick: every start state in the task curriculum points into the replay buffer $\mathcal{D}_{env}$, so its recorded continuation $\tau_\ell$ is already stored — auditing costs no new interaction. The gate rolls the world model forward from $\ell$ under the recorded actions, computes per-error terms $E_k$, standardizes them against fixed reference statistics (computed once on the backbone's random replay before training starts, never updated), and obtains the fidelity error (Eq. 6):
$$F_\theta(\ell)=\sum_{k\in C} w_k\frac{E_k(\tau_\ell)-\mu^{ref}_k}{\sigma^{ref}_k}$$
States with $F_\theta(\ell)>\epsilon_{WM}$ are quarantined: kept in $\mathcal{B}_G$ but never sampled, with their recorded continuations sent to $\mathcal{B}_{WM}$ for repair — directly cutting the feedback loop where a model error inflates learning potential and the policy keeps training on that error. Every $N_{gate}$ transitions the gate audits two groups: the top-$K$ highest-priority states in $\mathcal{B}_G$ (the ones the curriculum samples most, where model errors hurt most), and the full quarantine set (to check whether repair worked). After $n_{pass}$ consecutive passes a state's potential is recomputed under the repaired model and it returns to training; after $n_{fail}$ failures it is dropped from both curricula.
The threshold matters too: on SMACv2 $\epsilon_{WM}=0.5$ sits at the 72nd–83rd percentile of the backbone's fidelity error on its reference data — a fixed threshold, not a per-audit percentile, because a percentile would quarantine a fixed fraction of states no matter how accurate the model becomes; a fixed threshold automatically admits more states as the model improves. Set once, used across all nine scenarios and every method, zero tuning.
Experiments: The Gate Is the Whole Story
The paper evaluates on two families: nine procedurally-varying SMACv2 combat scenarios (Terran/Protoss/Zerg × 5v5/10v10/10v11, discrete control, MARIE actor–critic backbone) and the Multi-Agent Quadruped Environment (Gate-2/3 — pass a narrow gate without collision — and Shepherd — herd 1 or 9 sheep, continuous control, SeqWM planning backbone).

SMACv2: top row PROWL-2 vs the MARIE backbone, bottom row five model-free baselines (source: paper Figure 2)
SMACv2: PROWL-2 takes the highest mean win rate in all nine scenarios with the backbone architecture and objective untouched — the gap isolates the curriculum and repair mechanisms. The gap widens with scale: +4–18% over the backbone at 5v5, +20–91% at 10v10/10v11. Representative win rates (3 seeds):
| Scenario | Best model-free | MARIE backbone | PROWL-2 |
|---|---|---|---|
| Terran 5v5 | 21.2 (QPLEX) | 45.0 | 52.9 |
| Terran 10v11 | 3.3 (QMIX) | 5.6 | 10.7 |
| Protoss 5v5 | 16.1 (MAPPO) | 54.1 | 60.2 |
| Protoss 10v10 | 5.1 (MAPPO) | 27.9 | 34.8 |

MQE tasks: Gate-n — n quadrupeds pass a narrow gate as fast as possible without collision; Shepherd — two quadrupeds herd 1 (Easy) or 9 (Hard) sheep into a target region (source: paper Figure 3, renders from Xiong et al. 2024)

MQE success-rate curves over 1.5M environment steps (source: paper Figure 4)
MQE: matched difficulty pairs expose the differences. On Gate-2 and Shepherd-Easy all strong methods saturate (PROWL-2 99.2%, half a point from HASAC's 99.7%). The hard variants split the field:
| Task | All baselines | SeqWM backbone | Dual (ungated) | PROWL-2 |
|---|---|---|---|---|
| Gate-3 | 0.0 | 29.8 | 12.3 | 70.4 |
| Shepherd-Hard | 0.0 | 7.3 | 7.7 | 28.2 |
The ablations are the most informative part of the paper — the two curricula are not additive:
- Repair alone exceeds the backbone in all nine settings, but within one standard error in seven — the ticket was bought, the train was not boarded.
- The ungated curriculum (Dual) falls below the backbone in all nine SMACv2 scenarios — an ungated imagined curriculum is feeding the policy the model's hallucinations.
- Adding the gate turns the same curriculum into the largest single gain: Gate-3 jumps from 12.3% to 70.4%. PROWL-2's increment over Repair exceeds Repair's increment over the backbone in eight of nine settings.
Appendix ablations confirm the mechanism: quarantining the same number of states at random, or by model uncertainty (ensemble disagreement), does not work — only fidelity measured against real trajectories has discriminative power. Disagreement may serve as the developer's search heuristic, but never as the admission criterion.
Honest Boundaries
The paper lists its own limits: the 1M-step SMACv2 budget is five times MARIE's own, so the comparison covers both learning speed and final performance; the 1.5M-step MQE budget is below SeqWM's 4M (a compute constraint), both world-model learners are still improving on the hardest tasks, and whether the gap persists asymptotically is open. Fidelity targets (the choice of $E_k$) are hand-picked per domain. The developer's real-environment search costs 1–2× the task-policy interaction on SMACv2 (on MQE it is charged to the task budget, and PROWL-2 still leads). Single-agent settings are untested.
Why the Robotics Community Should Care
Odyssey positions PROWL-2 as one link in its "Experience Machines" agenda: as games, robotics and swarm systems increasingly rely on agents trained inside learned simulators, "which simulated experiences can be trusted" stops being an implementation detail and becomes a foundation-level question. The paper explicitly names humanoid-robot coordination — whose contact-rich dynamics make world-model errors more frequent — and long-horizon multi-agent games as next steps. Their foundation world model Odyssey-3 is described as the supplier of grounded imagined experience, with PROWL as the mechanism that keeps that experience trustworthy and progressively harder. The framework is fully agnostic to the world-model architecture and the MARL learner beneath it — it changes what is trained on, not how — so existing Dreamer-style or multi-agent world-model pipelines can adopt it as-is.
One-line summary: training agents inside world models requires deciding not only which imagined experience is useful but which is trustworthy — making that distinction explicit turns world-model failures from a cause of policy error into a source of policy improvement.
References
- Original post: Odyssey — Introducing PROWL-2: Jointly Learning Simulation and Decision-Making (Oct 1, 2026), odyssey.systems/introducing-prowl-2
- Paper PDF: prowl.odyssey.ml/PROWL-2.pdf (Güzel, Seidenschwarz, Hawke, Bogunovic — UCL AI Centre / Odyssey / University of Basel)
- Prior work: PROWL-1 (arXiv:2605.18803); IMAC (Güzel et al., NeurIPS 2025)
- Backbones: MARIE (Zhang et al., 2024, arXiv:2406.15836); SeqWM (Zhao et al., ICLR 2026)
- Benchmarks: SMACv2 (Ellis et al., 2023); MQE (Xiong et al., IROS 2024)
- Related: Odyssey's "The Era of Multi-Agent Imagined Experience", "Our Path to Superintelligence", Agora-2