GLM 5.3 DeepSWE v1.1: 66.9 Explained — What the Score Means
Quick answer: GLM 5.3 scores 66.9 on DeepSWE v1.1 — up from GLM 5.2's 46.2 and now within striking distance of the closed frontier (Fable 5: 69.7, GPT-5.6 Sol: 72.7). DeepSWE tests deep software engineering: complex, multi-file tasks that mimic real issue-to-PR work. It's one of the most workload-relevant benchmarks in Z.AI's launch table, and GLM 5.3's +20.7 jump is a headline post-training gain.
TL;DR
| Model | DeepSWE v1.1 |
|---|---|
| GLM 5.2 | 46.2 |
| GLM 5.3 | 66.9 (+20.7) |
| Kimi K3 | 67.5 |
| DeepSeek-V4 Pro | 62.7 |
| Qwen3.8-Max | 56.6 |
| Claude Opus 4.8 | 58.0 |
| Claude Fable 5 | 69.7 |
| GPT-5.6 Sol | 72.7 |
What Is DeepSWE?
DeepSWE (Deep Software Engineering) evaluates agents on complex, multi-file software engineering tasks — the kind that mimic real "issue → investigation → fix → PR" work rather than isolated coding puzzles. High scores require:
- Understanding a codebase across multiple files
- Diagnosing subtle bugs from partial information
- Making changes that don't break existing behavior
- Producing a complete, mergeable patch
It's one of the closest public benchmarks to what a working software engineer actually does — which is why the community watches it closely.
GLM 5.3's Jump: 46.2 → 66.9
The +20.7 gain is one of the largest in Z.AI's launch table (only Terminal-Bench 3.0's +23.7 is bigger). For context on the scale of the jump:
- GLM 5.2 → 5.3: 46.2 → 66.9 (+20.7) — a post-training-only improvement
- GLM 5.3 vs Kimi K3: 66.9 vs 67.5 — statistically a tie
- GLM 5.3 vs DeepSeek-V4 Pro: 66.9 vs 62.7 — GLM leads
- GLM 5.3 vs Opus 4.8: 66.9 vs 58.0 — GLM leads by ~9 points
The same base model, one month of scaled RL post-training, and DeepSWE jumps ~45%. That's the "post-training scaling works" story in one number.
Where It Still Trails
The closed frontier still holds the top of the DeepSWE table:
- Claude Fable 5: 69.7 (+2.8 over GLM 5.3)
- GPT-5.6 Sol: 72.7 (+5.8)
GLM 5.3 closes most of the gap but doesn't cross it — consistent with the overall pattern (GLM 5.3 approaches Fable 5 on coding, trails on the hardest engineering benchmarks).
Why DeepSWE Matters for You
- Most workload-relevant coding benchmark — issue-to-PR tasks ≈ real dev work.
- Predictive of agent quality — a 66.9 agent is meaningfully more useful in a real repo than a 46.2 one.
- The +20.7 jump is a post-training proof point — same base, much better engineering.
If you're evaluating GLM 5.3 for production coding agents, DeepSWE is the benchmark to quote: open-weights models within 3 points of Claude Fable 5 on real software engineering.
FAQ
What is GLM 5.3's DeepSWE v1.1 score? 66.9 — up from GLM 5.2's 46.2 (+20.7).
What does DeepSWE measure? Deep software engineering: complex multi-file tasks mimicking real issue-to-PR work.
How does GLM 5.3 compare on DeepSWE? Ties Kimi K3 (67.5), beats DeepSeek-V4 Pro (62.7) and Opus 4.8 (58.0), trails Fable 5 (69.7) and GPT-5.6 Sol (72.7).
Is DeepSWE a good benchmark? It's one of the most workload-relevant — real codebases, real engineering workflows — so scores transfer better than puzzle-style benchmarks.
Where do these numbers come from? Z.AI's official GLM-5.3 launch benchmark table (August 14, 2026).
Sources
- Z.AI: GLM-5.3: Frontier Coding with Emergent Cyber Capabilities (August 14, 2026)
- VentureBeat: GLM-5.3 launch coverage (August 14, 2026)
Last updated: August 18, 2026




