GLM 5.3 on AI Leaderboards: Where It Ranks (AA, Terminal-Bench, CyberGym)
Quick answer: GLM 5.3 holds three notable leaderboard positions: 8th of 181 models on Artificial Analysis' Intelligence Index (60) — tied with Kimi K3 and 3 behind Claude Opus 5 — plus open-weights SOTA on Terminal-Bench 3.0 (28.3) and the best public CyberGym result (84.5). The consistent theme: top-tier intelligence at open-weights pricing.
TL;DR
| Leaderboard | GLM 5.3 position |
|---|---|
| AA Intelligence Index | 60 — 8th of 181, tied Kimi K3 |
| Class median | 35 (GLM 5.3: 60, +25 over median) |
| Terminal-Bench 3.0 | 28.3 — open-weights SOTA |
| DeepSWE v1.1 | 66.9 — within 3 of Fable 5 |
| CyberGym | 84.5 — best public result |
| AutomationBench | 48.2 — best in Z.AI's official table |
| Cost per AA task | $0.68 — lowest among top scorers |
Artificial Analysis Intelligence Index: 60
The independent Artificial Analysis evaluation (published August 18, 2026) placed GLM 5.3 at 60 on the Intelligence Index at max reasoning effort:
- 8th of 181 models in its class (median: 35)
- Tied with Kimi K3 (60) — the top-scoring open-weights model
- 3 behind Claude Opus 5 (63, leader)
- Composite of 9 evals: agentic work, tool use, terminal coding, scientific reasoning, graduate science, physics, knowledge reliability, long-context
The cost angle is where GLM 5.3 stands out: $0.68 per index task vs $0.84 (Kimi K3) and $2.34 (Opus 5) — cheapest top-10 model per unit of measured intelligence.
Terminal-Bench 3.0: Open-Weights SOTA
Z.AI's official table shows GLM 5.3 at 28.3 on Terminal-Bench 3.0 — the hardest public terminal benchmark:
- GLM 5.2: 4.6 → GLM 5.3: 28.3 (+23.7)
- Open-weights SOTA, ahead of Kimi K3 (17.4)
- Close to closed frontier: Fable 5 (33.7), GPT-5.6 Sol (34.6)
CyberGym: Best Public Result
The surprise leaderboard: CyberGym 84.5 — best public vulnerability-discovery result:
- Ahead of GPT-5.6 Sol (83.6) and Mythos 5 (83.8)
- ExploitBench 54.4 (more than double GLM 5.2's 24.4)
- ExploitGym 105/130 tasks (vs 29/39 for GLM 5.2)
AutomationBench: Best in Table
AutomationBench v1.0.6: 48.2 — GLM 5.3 leads the entire official comparison (Kimi K3: 46.7, GPT-5.6 Sol: 45.8) on autonomous business-process automation.
How to Read These Rankings
- AA is independent; Z.AI's table is vendor-reported. Both agree GLM 5.3 is top-tier open-weights.
- Workload matters. GLM 5.3 leads on terminal/automation/security; trails on FrontierSWE/ProgramBench (Fable 5, GPT-5.6 Sol lead).
- Price changes the calculus. Score-60 at $0.68/task and open weights is a different proposition than score-63 at $2.34/task closed.
FAQ
Where does GLM 5.3 rank on leaderboards? AA Intelligence Index 60 (8th of 181), Terminal-Bench 3.0 open SOTA (28.3), CyberGym best public (84.5), AutomationBench best in table (48.2).
Is GLM 5.3 the best open-weights model? On AA, it ties Kimi K3 for the top open-weights score (60). On Terminal-Bench 3.0 and CyberGym, it's the open leader.
Is GLM 5.3 better than Claude Opus 5? AA: Opus 5 leads 63 vs 60. GLM 5.3 wins on cost ($0.68 vs $2.34 per task) and openness.
What is GLM 5.3's Artificial Analysis score? 60 on the Intelligence Index (max reasoning effort), published August 18, 2026.
Where do these rankings come from? Artificial Analysis (independent) and Z.AI's official launch benchmark table.
Sources
- Artificial Analysis: GLM-5.3
- Unite.AI: GLM-5.3 Scores 60 on Artificial Analysis Intelligence Index (August 18, 2026)
- Z.AI: GLM-5.3: Frontier Coding with Emergent Cyber Capabilities (August 14, 2026)
Last updated: August 18, 2026




