Skip to content
← Tags

#Fidelity-Gate (1)

Introducing PROWL-2: Fidelity-Gated Dual Curricula — Jointly Learning Simulation and Decision-Making

Introducing PROWL-2: Fidelity-Gated Dual Curricula — Jointly Learning Simulation and Decision-Making

Odyssey (with UCL AI Centre and the University of Basel) released PROWL-2 on Oct 1: the first framework coupling an agent curriculum and a world-model repair curriculum inside one continual training loop. Core insight: in imagination training, a high-learning-signal trajectory is ambiguous — a real policy weakness or just a wrong world-model prediction; prioritizing it indiscriminately reinforces the model's own hallucinations. A fidelity gate separates 'useful' from 'trustworthy': reliable high-potential imagined experience feeds the policy curriculum, unreliable rollouts — with their already-stored real continuations — go to a repair pool, fixed by a randomly-initialized, KL-anchored developer policy exploring the real environment, re-audited and readmitted once repaired. First place on all nine SMACv2 scenarios (+4-18% at 5v5, +20-91% at 10v10/10v11 over the backbone); big margins on hard MQE tasks — Gate-3 70.4% vs 29.8%, Shepherd-Hard 28.2% vs 7.3%, all other baselines at 0. Ablations show the curricula are not additive: repair alone is marginal, the ungated curriculum falls below the backbone in all nine scenarios, and the gate turns the same curriculum into the largest single gain. Architecture- and algorithm-agnostic; next steps point at humanoid coordination and long-horizon multi-agent games.

BLOG

PROWL-2World ModelsMARLReinforcement LearningCurriculum Learning