General reasoning, scientific QA, long-chain inference and planning
Nothing tracked yet
SOTA RADAR
Thirteen capability domains in four families. Each cell answers three things: what ruler measures this domain, who sits on top of the ladder right now, and how far that claim can be trusted. We do not keep a full tool directory — the long tail stays behind the inclusion gate, and only the ladder shows.
General reasoning, scientific QA, long-chain inference and planning
Nothing tracked yet
Code generation, repair, refactoring, software engineering tasks
Nothing tracked yet
Competition math, proof assistance, numerical reasoning
Nothing tracked yet
Image-text understanding, chart/doc QA, visual reasoning
Nothing tracked yet
Text-to-image, image editing, consistency and controllability
Nothing tracked yet
Text/image-to-video, physical and temporal coherence
Nothing tracked yet
3D asset generation, scene reconstruction, long-horizon world prediction
Nothing tracked yet
TTS naturalness, voice cloning, music generation, ASR
Nothing tracked yet
Multi-step tool use, web/OS tasks, long-horizon autonomy
Nothing tracked yet
Evaluation infra, reproducibility, failure taxonomy, cost & latency
Nothing tracked yet
Long-horizon policy in resettable environments, open-world game agents
Nothing tracked yet
Real-robot manipulation, navigation, VLA and whole-body control
Nothing tracked yet
Jailbreak resistance, hallucination rate, tool misuse, long-task boundaries
Nothing tracked yet
Applied only to SOTA claims and benchmark evidence. The letter on the seal says how far a claim is from something we reproduced ourselves — C and D are deliberately drawn dashed and grey so they never read as hard as an A.