Skip to content
←Back to the radar

SOTA

LLM

The foundation-model layer: the weights themselves plus what is tightly coupled to them - training recipes (pretraining, post-training, RL/RLHF), architecture and long context, and the inference serving stack with its cost per token

RULERIntelligence: Artificial Analysis Intelligence Index, GPQA, MMLU-Pro, Arena Elo. Capacity: effective long-context recall, modalities. Supply: tok/s, $ per 1M tokens, open-weight licence

The ladder

6
ReasoningDevelopmentTopA

Claude Opus 5.5: the long-running agentic tier at $4/$20 with 1M context and adaptive thinking

The fixed version Anthropic positions for long-running agentic coding and knowledge work (claude-opus-5-5, retirement no earlier than 2027-09-22): 1M-token context, 128K-token single output, adaptive thinking defaulting to medium, at $4/$20 per million tokens, the cheapest of the three top rows we track. On the Artificial Analysis intelligence index this site syncs (2026-09-22) the top three rows are all its thinking tiers (max 57.62, xhigh 55.99, high 53.58), while the default medium tier sits #8 at 51.24, with sibling Fable 5.1 (53.35) and GPT-6 Astra (52.67) in between. Two discounts: those rows carry the with fallback qualifier while Astra reads (max), so the protocols differ, and the 53.94% at #8 on Terminal-Bench belongs to the previous-generation Opus 5. Closed, API only; adaptive thinking is a black box, so budget on p90 rather than the mean; 1M context is not 1M of effective attention, so whole-repo input still needs retrieval. Use it for most workloads and move to Fable 5.1 only when the highest tier is not enough: 2.5x the price for 4.3 index points. Graded A (confirmed); not benchmarked by us.

57.62AA 智能指数Confirmed · 2026-09
ProductionAnthropicSite
Claude Opus 5.5: the long-running agentic tier at $4/$20 with 1M context and adaptive thinking
ReasoningDevelopmentTopC

Jev: TypeSafe's first System One Model types decisions and prices them two orders of magnitude lower

TypeSafe's first System One Model: outputs are type-safe structured values under three question primitives (Noul/Choice/Score) rather than strings, trained with RLCD for calibrated decision probabilities. The official 13-question parallel experiment makes one synthesized call ~12.2x cheaper and ~10.0x faster than 13 separate calls; four-workflow mean accuracy 67.8% at /tmp/run_ingest2.sh.0004 and 0.4s per instance - opus 5 / sol tier accuracy at almost two orders of magnitude less cost; input $0.042/M tokens, output free. Ships an OpenJev local reproduction path and an eight-item caveats list.

67.8%四工作流平均准确率(官方评测)Vendor Claim · 2026-09
ProductTypeSafeSiteRepo
Jev: TypeSafe's first System One Model types decisions and prices them two orders of magnitude lower
ReasoningDevelopmentTopA

MiMo-V2.6: Live RL as an auditable ledger of 30 steps, 750k trajectories, and dollars

Xiaomi's open-source MiMo-V2.6 Pro/Flash: 30 steps of Live RL, roughly 750k trajectories, public cost of $850k/$2.62M, and 7k+ RL environments released in the same batch; tops the open-weights camp on the Artificial Analysis intelligence index at 46 with +17/+14 points out-of-sample on DeepSWE v1.1; four demo lines (Vibe World, CUA, research, content creation) all open. Self-improvement becomes an auditable ledger - steps, trajectories, and dollars on the table for third-party recomputation.

46AA 智能指数(开源榜首)Confirmed · 2026-09
ProductXiaomiSiteRepo
MiMo-V2.6: Live RL as an auditable ledger of 30 steps, 750k trajectories, and dollars
ReasoningDevelopmentTopA

Claude Fable 5.1: the long-horizon agentic tier that tops the ladder on a different ruler

The fixed version Anthropic positions for demanding reasoning and long-horizon agentic work (claude-fable-5-1, retirement no earlier than 2027-09-01): default thinking effort high at $10/$50 per million tokens, 2.5x sibling Opus 5.5, and default-high is itself a default cost behaviour. Its standing only holds once you switch rulers: #1 on net_improvement in the Arena Agent board (13.71, confirmed_success 19.83, praise 31.83) ahead of GPT-6 Astra (11.54/17.70/32.79), while Opus 5.5 misses the top eight; #2 on Terminal-Bench 4.0 at 57.88% (n=330), 0.3 points behind Astra at 58.18% but with pass@5 of 0.7879 against 0.7121, lower single-shot and more robust over five attempts; #4 on the intelligence index at 53.35 (max), below Opus 5.5. The picture is consistent: not first on composite intelligence, first on pushing a real task forward. A separate praise column means the board carries a human or judge component and is not an objective benchmark. Choose by whether the workload is hard from reasoning or from long-horizon consistency. Closed, API only. Graded A (confirmed); no long-task comparison on our own harness.

13.71Arena Agent 净改进Confirmed · 2026-09
ProductAnthropicSite
Claude Fable 5.1: the long-horizon agentic tier that tops the ladder on a different ruler
ReasoningDevelopmentTopA

GPT-6 Astra: first on terminal engineering tasks, third tier on composite intelligence

The OpenAI flagship tier, officially positioned for hardest end-to-end work (gpt-6-astra): 1.05M-token context, 128K-token single output, knowledge cutoff 2026-04-30, with Functions, Web search, File search and Computer use built in, so search, retrieval and interface control ship with the model instead of an outer agent framework. Reasoning effort runs low to max in five steps while price stays fixed at $10/$50, so the cost lever is token consumption; siblings Sol ($2/$10) and Luna ($0.1/$0.5) make a 100x spread, so load splitting can stay inside one vendor. Readings (2026-09-22): #1 on Terminal-Bench 4.0 at 58.18% (n_trials=330, pass@5 0.7121), the protocol closest to a coding agent's daily life; #2 on Arena Agent with net_improvement 11.54 and confirmed_success 17.70, both below leader Fable 5.1, but the field's highest praise at 32.79; #6 and #7 on the intelligence index (max 52.67). The cutoff is nearly five months older than this page, so version numbers and API changes must go through Web search; computer-use reliability appears on no board; closed, API only. Graded A (confirmed); not benchmarked by us.

58.18%Terminal-Bench 准确率Confirmed · 2026-09
ProductOpenAISite
GPT-6 Astra: first on terminal engineering tasks, third tier on composite intelligence
ReasoningDevelopmentC

Laya: a 421M non-autoregressive System 1 decision model at 33 ms per forward pass

Convai Innovations' open-source non-autoregressive System 1 decision model: 421M parameters, ModernBERT backbone with [MASK] option extraction, one 33 ms forward pass per calibrated decision, Apache-2.0; trained with RLCD strictly proper scoring rules for honest probabilities; the model card ships a real benchmark against Jev plus an honest limitations list. Positioned for edge and high-concurrency small decisions (game NPCs, dialogue policy, request routing), not long reasoning or open generation.

33 ms单次前向延迟(421M,官方)Vendor Claim · 2026-09
ResearchConvai InnovationsSiteRepo
Laya: a 421M non-autoregressive System 1 decision model at 33 ms per forward pass

Evidence

9
2026-09-12Sam Altman: The Next ChatGPT Will Watch Your Screen and Remember EverythingSam Altman says that within about six months the next generation of ChatGPT will watch your screen, listen to your meetings and calls, and quietly assemble a persistent context of your work. The real watershed, he argues, is no longer compute but memory.2026-07-24$8 ESP32-S3 Runs 29M-Parameter LLMA developer forces a 29M-parameter LLM onto an $8 ESP32-S3 MCU, running fully offline at LED-level power draw.2025-08-27ReST-RL: Reinforcing LLM Reasoning through Unified Self-Training and Value-Guided SearchGRPO is the representative RL method for improving LLM reasoning, yet it only ever sees one sparse reward at the end of a whole trajectory: when the rewards inside a sampled group land close together, the group-relative advantage collapses into noise and the policy learns almost nothing. ReST-RL reconnects policy optimization and value-guided search into a single self-training pipeline. Stage one, ReST-GRPO, first filters out low-information prompts by reward standard deviation, then draws prefixes from each prompt's highest-reward trajectory under a discrete exponential distribution and uses them as fresh online-GRPO starting contexts. A selected prefix is context only; its suffix is re-sampled and optimized rather than imitated. Stage two, VM-MCTS, runs MCTS under the now-static policy to self-collect value targets and trains a value model that predicts expected terminal reward. At inference the same model both allocates the tree search through UCT and ranks completed candidates in a Best-of-N fashion, so search and verification share one state-value scale. On coding benchmarks including APPS, BigCodeBench and HumanEval, Qwen3-8B moves from 0.503 to 0.689 average. In matched policy-value controls, ReST-GRPO + VM-MCTS reaches 0.642 on APPS-500 while GRPO + VM-MCTS reaches only 0.538, so the stage-one distributional shift survives value learning. End-to-end accounting puts ReST-GRPO at 1,752 GPU-hours against 2,080 for GRPO, hitting a 9% gain in 71 hours instead of 207. A value model trained only on code trajectories also transfers to MATH, Omni-MATH and GPQA-Diamond without target-domain tuning.

How it is used

2026-09-22MiMo-V2.6 Deep Read: Six Days of Live RL, an AA Index of 46 and a Fully Open Self-Improvement RunXiaomi open-sources MiMo-V2.6 Pro/Flash: 30 live RL steps, ~750k trajectories at $850k/$2.62M; AA Index 46 tops open models, DeepSWE v1.1 gains +17/+14 out of sample; Vibe World, CUA, science and content demos plus 7k+ RL environments released.2026-09-21Laya, the Open-Source System 1 Decision Model: 421M Parameters, One 33 ms Forward PassConvai Innovations' open answer to TypeSafe Jev: a 421M non-autoregressive decision model trained with RLCD proper scoring rules, plus the README's honest limits and real Jev comparisons.2026-09-20Build Your Own Jev (100% Local): Recreating Fixed-Answer Scoring with SGLangJev-style fixed-answer scoring reproduced locally with SGLang /v1/score: single-token labels, restricted softmax, confidence-gated routing, and a 6 ms vs 1088 ms benchmark against generation.2026-09-18Jev and System One Models: Frontier Intelligence as a Function CallA deep read of TypeSafe's first System One Model: three question primitives, the economics of parallel calls, confidence-gated routing, eval caveats, eight jagged edges, and the OpenJev local repro.2026-09-01Fine-tune, Convert and Deploy Language Models on NVIDIA Jetson with Unsloth and llama.cppOfficial Jetson AI Lab tutorial: fine-tune models directly on Jetson with Unsloth and JetPack 7.2 using memory-efficient QLoRA, export to GGUF, and run locally with llama.cpp. Two hands-on examples: Qwen3.5-4B vision-language model on Jetson Orin Nano (LaTeX OCR fine-tuning), and NVIDIA Nemotron 3.5 Lightning 30B-A3B on Jetson AGX Thor (3 training steps in 44.8s, 66.1 tokens/sec with Q4_K_M quantization). No cloud required.2025-08-29Inside vLLM: Anatomy of a High-Throughput LLM Inference SystemA systems walkthrough of vLLM: engine loop, scheduler, paged-attention KV blocks, continuous batching, prefix caching, speculative decoding, disaggregated P/D, multi-GPU executors, a two-node serving stack, and the latency-vs-throughput roofline.

Assets you can use now

No installable skill wired to this domain yet.

Boundaries & failure modes

The boundary note for this domain is not written yet. It should answer three things: where it is stuck, under which conditions it breaks, and the order of magnitude of cost and latency. Until it lands, a ladder position only says this asset leads on our ruler — not that it works on your task.