GLM 5.3 on AI Leaderboards: Where It Ranks (AA, Terminal-Bench, CyberGym)

GLM 5.3 on AI Leaderboards: Where It Ranks (AA, Terminal-Bench, CyberGym)

GLM 5.3 leaderboard positions — AA Intelligence Index 60 (8th of 181, tied Kimi K3), Terminal-Bench 3.0 open SOTA 28.3, CyberGym 84.5 best public. Full ranking context.

GLM 5.3 on AI Leaderboards: Where It Ranks (AA, Terminal-Bench, CyberGym)

Quick answer: GLM 5.3 holds three notable leaderboard positions: 8th of 181 models on Artificial Analysis' Intelligence Index (60) — tied with Kimi K3 and 3 behind Claude Opus 5 — plus open-weights SOTA on Terminal-Bench 3.0 (28.3) and the best public CyberGym result (84.5). The consistent theme: top-tier intelligence at open-weights pricing.


TL;DR

LeaderboardGLM 5.3 position
AA Intelligence Index60 — 8th of 181, tied Kimi K3
Class median35 (GLM 5.3: 60, +25 over median)
Terminal-Bench 3.028.3 — open-weights SOTA
DeepSWE v1.166.9 — within 3 of Fable 5
CyberGym84.5 — best public result
AutomationBench48.2 — best in Z.AI's official table
Cost per AA task$0.68 — lowest among top scorers

Artificial Analysis Intelligence Index: 60

The independent Artificial Analysis evaluation (published August 18, 2026) placed GLM 5.3 at 60 on the Intelligence Index at max reasoning effort:

  • 8th of 181 models in its class (median: 35)
  • Tied with Kimi K3 (60) — the top-scoring open-weights model
  • 3 behind Claude Opus 5 (63, leader)
  • Composite of 9 evals: agentic work, tool use, terminal coding, scientific reasoning, graduate science, physics, knowledge reliability, long-context

The cost angle is where GLM 5.3 stands out: $0.68 per index task vs $0.84 (Kimi K3) and $2.34 (Opus 5) — cheapest top-10 model per unit of measured intelligence.

Terminal-Bench 3.0: Open-Weights SOTA

Z.AI's official table shows GLM 5.3 at 28.3 on Terminal-Bench 3.0 — the hardest public terminal benchmark:

  • GLM 5.2: 4.6 → GLM 5.3: 28.3 (+23.7)
  • Open-weights SOTA, ahead of Kimi K3 (17.4)
  • Close to closed frontier: Fable 5 (33.7), GPT-5.6 Sol (34.6)

CyberGym: Best Public Result

The surprise leaderboard: CyberGym 84.5 — best public vulnerability-discovery result:

  • Ahead of GPT-5.6 Sol (83.6) and Mythos 5 (83.8)
  • ExploitBench 54.4 (more than double GLM 5.2's 24.4)
  • ExploitGym 105/130 tasks (vs 29/39 for GLM 5.2)

AutomationBench: Best in Table

AutomationBench v1.0.6: 48.2 — GLM 5.3 leads the entire official comparison (Kimi K3: 46.7, GPT-5.6 Sol: 45.8) on autonomous business-process automation.

How to Read These Rankings

  1. AA is independent; Z.AI's table is vendor-reported. Both agree GLM 5.3 is top-tier open-weights.
  2. Workload matters. GLM 5.3 leads on terminal/automation/security; trails on FrontierSWE/ProgramBench (Fable 5, GPT-5.6 Sol lead).
  3. Price changes the calculus. Score-60 at $0.68/task and open weights is a different proposition than score-63 at $2.34/task closed.

FAQ

Where does GLM 5.3 rank on leaderboards? AA Intelligence Index 60 (8th of 181), Terminal-Bench 3.0 open SOTA (28.3), CyberGym best public (84.5), AutomationBench best in table (48.2).

Is GLM 5.3 the best open-weights model? On AA, it ties Kimi K3 for the top open-weights score (60). On Terminal-Bench 3.0 and CyberGym, it's the open leader.

Is GLM 5.3 better than Claude Opus 5? AA: Opus 5 leads 63 vs 60. GLM 5.3 wins on cost ($0.68 vs $2.34 per task) and openness.

What is GLM 5.3's Artificial Analysis score? 60 on the Intelligence Index (max reasoning effort), published August 18, 2026.

Where do these rankings come from? Artificial Analysis (independent) and Z.AI's official launch benchmark table.


Sources

Last updated: August 18, 2026

Start Using GLM 5 Today

Try GLM 5 free — reasoning, coding, agents, and image generation in one platform.