GLM-5.3: the open-weights coding flagship bought with post-training alone
glm-5-3
Z.ai post-trained the very same base as GLM-5.2 and nothing else, producing the 744B-A40B GLM-5.3: 50 percent over 5.2 on the in-house Z.ai Code Bench, open-weights first on Terminal Bench 3.0 at 28.3 and Agents' Last Exam at 28.5, while AutomationBench 48.2 and GDPVal-AA v2 1769 beat every closed-source comparison on the official chart; CyberGym vulnerability discovery is SOTA and the exploitation-chain gains emerged without dedicated training. The companion 320B-A18B Flash starts from a new base and is the first GLM to mix sparse with linear attention, attacking long-context serving cost. Weights ship in FP8 and BF16, deployable on SGLang, vLLM, KTransformers and the Ascend stack.
- CONFIDENCE
- Vendor Claim
- Official model card or keynote only, no independent re-test
- KEY METRIC
- Terminal Bench 3.0(开源权重第一, 官方图)
- Vendor Claim · 2026-08
- MATURITY
- Product
- research → demo → product → production
Our takeWaterline: same base, post-training only, and Terminal Bench 3.0 moves from 4.6 to 28.3 while AutomationBench moves from 26.2 to 48.2, a generational slope that is itself the strongest evidence yet for open-weights post-training recipes. All six readings come from the official chart, DeepSWE still goes to Kimi K3 by 0.6, and no third-party reproduction is public, so confidence is C (vendor-claim). Against the closed frontier it stays 5-6 points behind on Terminal Bench 3.0; long-horizon agentic and automation benchmarks are the fronts where it draws or leads. Pick 5.3 for the strongest open-weights coding and long-horizon agent, Flash for cost per token; keep reasoning_effort at its default max when reproducing the boards and pass clear_thinking=true explicitly in chat.
In one line: the open-weights coding flagship bought entirely with post-training
GLM-5.3 is a 744B-A40B MoE that Z.ai (Zhipu) obtained by post-training the same base model as GLM-5.2, with weights published in FP8 and BF16 on HuggingFace and ModelScope. The companion release GLM-5.3-Flash is a 320B-A18B model on a newly trained base. The official claim is a single sentence: the most capable open-weights model for coding, 50 percent better than 5.2 on the in-house Z.ai Code Bench, and open-source SOTA on Terminal Bench 3.0 and Agents' Last Exam.
Same base, post-training only is the most information-dense fact in this release: not one byte of pre-training changed, every gain came from the post-training recipe and compute. It turns how much headroom post-training still holds into a measurable question, and it shows every open team chasing the closed frontier a route that does not require re-training a base model.
The data surface: two weight tiers, two bases, one deployment matrix
| Weights | Size | Precision | Source and notes |
|---|---|---|---|
| GLM-5.3 | 744B-A40B | FP8 | HuggingFace zai-org/GLM-5.3, ModelScope ZhipuAI/GLM-5.3 |
| GLM-5.3-BF16 | 744B-A40B | BF16 | Same, for teams that quantize themselves |
| GLM-5.3-Flash | 320B-A18B | FP8 | New base, sparse plus linear hybrid attention, mHC |
| GLM-5.3-Flash-BF16 | 320B-A18B | BF16 | Same |
The deployment matrix covers SGLang, vLLM, Transformers (glm_moe_dsa and glm5_next), KTransformers and Unsloth; Flash additionally lists TokenSpeed plus the Ascend side with vLLM-Ascend and xLLM. Fine-tuning goes through slime v0.3.0+ or ms-swift v4.4.0+. The FP8 weights of a 744B model do not fit one machine, but 40B active parameters keep the throughput arithmetic payable, which is exactly the intent DSA sparse attention has carried all along: bend long-context cost into a shape that can actually be served.
Reading the six benchmarks: two were claimed, four were beaten
| Benchmark | GLM-5.3 | GLM-5.2 | Kimi K3 | Fable 5 | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal Bench 3.0 | 28.3 | 4.6 | 17.4 | 33.7 | 34.6 |
| DeepSWE | 66.9 | 46.2 | 67.5 | 69.7 | 72.7 |
| Agents' Last Exam (CLI) | 28.5 | 23.8 | 27.6 | 23.8 | 28.6 |
| AutomationBench | 48.2 | 26.2 | 46.7 | 46.2 | 45.8 |
| HLE w/ Tools | 62.5 | 54.7 | 59.8 | 63.9 | 64.5 |
| GDPVal-AA v2 | 1769 | 1508 | 1682 | 1743 | 1730 |
Four readings. First, the open-source SOTA claim in the README circles only Terminal Bench 3.0 and Agents' Last Exam, and on those two the chart does lead (28.3 against 17.4 for Kimi K3, 28.5 against 27.6), so claim and chart agree. Second, on DeepSWE the 67.5 of Kimi K3 edges out 66.9, so open-source first must not be extrapolated to all six rows; this site writes what the chart shows and stops there. Third, AutomationBench 48.2 and GDPVal-AA v2 1769 beat every closed-source comparison on the chart (Fable 5 and GPT-5.6 Sol), which is beyond the official claim and worth recording on its own. Fourth, the generational gap over GLM-5.2 is an order of magnitude: Terminal Bench 3.0 from 4.6 to 28.3, AutomationBench from 26.2 to 48.2. That slope, from post-training alone on an unchanged base, is the real news.
Flash: the tier that attacks the supply-side ruler
| Benchmark | GLM-5.3-Flash | GLM-5.2 | DeepSeek-V4-Vision-Exp | Claude Opus 4.8 | GPT-5.6 Terra | Gemini 3.7 Flash |
|---|---|---|---|---|---|---|
| Terminal-Bench 2.1 | 84.3 | 81.0 | 83.9 | 85.0 | 87.4 | 85.8 |
| DeepSWE v1.1 | 63.4 | 46.2 | 59.3 | 58.0 | 69.6 | 65.3 |
| Agents' Last Exam | 26.3 | 20.4 | 27.3 | 27.0 | 28.0 | - |
| AutomationBench v1.0.6 | 48.8 | 26.2 | 38.8 | 41.0 | 37.2 | 52.3 |
| HLE w/ Tools | 55.3 | 54.7 | 55.1 | 57.9 | - | - |
| GDPVal-AA v2 | 1773 | 1504 | 1675 | 1582 | 1571 | 1527 |
The point of Flash is not topping the chart but the architecture: for the first time in the GLM series, sparse attention and linear attention live in one hybrid architecture, sharply cutting long-context serving cost while keeping precise long-context capability; mHC (Manifold-Constrained Hyper-Connections) lifts scaling efficiency on top of a 30T-token multimodal pre-training corpus. On the chart it takes Terminal-Bench 2.1 84.3, AutomationBench 48.8 and GDPVal-AA v2 1773, leading GLM-5.2 across the board and trading wins with the closed frontier. It answers a different question: for the same intelligence, how much compute and VRAM per token.
Lineage: three base changes, two ruler changes
| Version | Size | Pre-training | Signature step |
|---|---|---|---|
| GLM-4.5 | 355B-A32B | 23T | The MoE foundation |
| GLM-5 | 744B-A40B | 28.5T | DSA plus slime async RL; open-source first on Vending Bench 2 (4,432 dollars) |
| GLM-5.1 | Same base as 5 | - | SWE-Bench Pro SOTA; leads NL2Repo and Terminal-Bench 2.0 |
| GLM-5.2 | Same base as 5 | - | Solid 1M context (IndexShare); TB2.1 81.0, SWE-bench Pro 62.1 |
| GLM-5.3 | 744B-A40B | Same base as 5.2 | Pure post-training; Terminal Bench 3.0 28.3, open-weights first |
| GLM-5.3-Flash | 320B-A18B | 30T multimodal | Sparse plus linear hybrid attention, mHC, attacks serving cost |
The cadence reads cleanly: 4.5 lays the foundation, 5 changes the base and adds DSA with async RL, 5.1 takes SWE-Bench Pro SOTA, 5.2 makes 1M context solid (IndexShare reuses the indexer every 4 layers, cutting per-token FLOPs by 2.9x at 1M and lifting MTP speculative accept length by 20 percent), 5.3 closes the coding and long-horizon gap on the same base with post-training alone, and Flash changes the base once more to attack supply. Technical report at arXiv 2602.15763; IndexShare at arXiv 2603.12201.
Deploying and reproducing: three traps that bite
First, reasoning_effort accepts low, high and max, and defaults to max when omitted (any other value also falls back to max); reproducing the boards requires keeping the default max, while GLM-5.2 accepts only high and max, so cross-version scripts need a branch. Second, in the chat template clear_thinking defaults to false when omitted, so a multi-turn conversation that does not pass true explicitly carries the previous turn thinking chain into context, changing both the token bill and the behavior. Third, weights and inference stack must be paired: 5.3 runs on glm_moe_dsa, 5.3-Flash on glm5_next, and running new weights under an old configuration tends to fail silently rather than raise.
Emergent cyber capability: why this is also evidence for the safety domain
The vendor itself admits cyber capability developed faster than expected: as post-training scaled, GLM-5.3 reached SOTA on CyberGym for vulnerability discovery, with the largest gains further up the exploitation chain, more than doubling GLM-5.2 on exploitation benchmarks, on capability that was never trained separately. Emergent capability cannot be predicted from a training manifest, so it cannot be recalled by one either; the attack surface of open weights grows with post-training scale, and evaluation plus disclosure cadence has to keep up. This release therefore also feeds the evidence stream of the safety domain.