GLM 5.3 Flash Benchmarks: Vendor Claims vs. Independent Numbers
Aug 27, 2026

GLM 5.3 Flash Benchmarks: Vendor Claims vs. Independent Numbers

GLM 5.3 Flash benchmarks decoded: 84.3 on Terminal-Bench 2.1, 63.4 DeepSWE, 57 on Artificial Analysis. Which numbers are first-party and which survive scrutiny.

Quick answer: Z.ai reports GLM 5.3 Flash at 84.3 on Terminal-Bench 2.1 (vs Claude Opus 4.8's 85.0), 63.4 on DeepSWE v1.1 (vs GLM-5.2's 46.2), 48.8 on AutomationBench (vs 26.2), 29.0 on its own Z.ai Code Bench v1.0 (vs Opus 4.8's 29.5), 62.4 on OfficeQA Pro, and 55.3 on HLE with tools. Independently, Artificial Analysis scores it 57 on the Intelligence Index at 50.2 output tokens/second with a 1.47s TTFT. The headline is real; the caveat is that most of the flattering comparisons are vendor-selected.

Benchmark tables are the most-copied and least-interrogated artefact in AI launches. A vendor publishes six numbers, forty blog posts reprint them, and nobody notes which benchmark the vendor authored, which comparison models were chosen, or what "max effort" means in the footnote. This page separates what Z.ai claims from what an independent lab measured, and tells you which of the two should move your decision.

Every figure below traces to either Z.ai's own model card and repository, or to Artificial Analysis's independent evaluation, checked August 27, 2026. We have not re-run these benchmarks ourselves — running Terminal-Bench properly is a serious compute commitment — and we would rather tell you that than imply testing we did not do. What we have done is check every number against its primary source rather than against another blog.

What This Article Solves

The pain point: a benchmark table tells you nothing until you know who wrote the benchmark and who picked the comparison. GLM 5.3 Flash's launch table includes one benchmark authored by Z.ai itself, several with flattering-but-narrow comparison sets, and one metric ("max effort") that is not directly comparable to how you would run the model in production.

You will leave knowing which three numbers are load-bearing, which one is marketing, what the independent lab says, and — the piece nobody publishes — the benchmark result that cuts against the model.

The First-Party Table

Z.ai's reported results, with the comparison models Z.ai selected:

BenchmarkGLM 5.3 FlashComparison
Terminal-Bench 2.184.3Claude Opus 4.8: 85.0 · GPT-5.6 Terra: 87.4
DeepSWE v1.163.4GLM-5.2: 46.2
AutomationBench48.8GLM-5.2: 26.2
Z.ai Code Bench v1.0 (max effort)29.0Claude Opus 4.8: 29.5
OfficeQA Pro62.4reported ahead of Opus 4.8
HLE (with tools)55.3

Now the grading.

Terminal-Bench 2.1 — 84.3. Load-bearing. This is a third-party agentic benchmark that measures whether a model can actually complete multi-step terminal tasks, and it is one of the hardest things to game. Landing 0.7 points behind Claude Opus 4.8 at a fraction of the price is the single most impressive claim in the launch. The honest asterisk: GPT-5.6 Terra sits at 87.4, so this is "near the frontier," not "at the frontier."

DeepSWE v1.1 — 63.4 vs GLM-5.2's 46.2. Load-bearing. This is a within-family comparison, same evaluation harness, same team, and a 37% relative improvement. Vendor-run, yes — but a vendor comparing against its own previous model has far less incentive to cook the methodology than one comparing against a competitor. Combined with the ~10x price reduction Z.ai claims over GLM-5.2, this is the number that justifies a migration.

AutomationBench — 48.8 vs 26.2. Load-bearing, with a caveat. An 86% relative jump is enormous, and enormous jumps usually mean the previous model was badly suited to the task rather than the new one being twice as smart. Read it as "GLM-5.2 was weak here and Flash fixed it," not as a general intelligence multiplier.

Z.ai Code Bench v1.0 (max effort) — 29.0 vs Opus 4.8's 29.5. Marketing. Two problems. First, Z.ai wrote this benchmark, and a vendor's in-house eval is not evidence about a competitor. Second, "max effort" is a reasoning-budget setting — the model is being run at its most expensive configuration to produce this number, which is not the configuration your cost model assumes. Discount accordingly.

OfficeQA Pro — 62.4. Interesting. This is the only benchmark where the Flash tier claims an outright win over a flagship. It is a document-and-visual reasoning benchmark, and it lines up with the model's genuine structural advantage: GLM 5.3 Flash is natively multimodal, trained on a 30T-token multimodal corpus, rather than a text model with a vision adapter. If your workload is documents, screenshots, and diagrams, this is the number to test against your own data first.

HLE with tools — 55.3. Context-free. Published without a comparison set on the model card, so it establishes a level but not a ranking.

The Independent Numbers

Artificial Analysis, which runs its own standardised evaluations rather than reprinting vendor tables, reports:

MetricGLM 5.3 FlashClass context
Intelligence Index57median ~27 for open-weight models of similar size
Output speed50.2 tok/smedian 67.0 t/s — below class median
Time to first token1.47 smedian 2.14 s — better than class median
Blended price$0.10 / 1M7:2:1 cache:input:output mix
Context window1.0M

The Intelligence Index of 57 corroborates the broad story: this model performs roughly twice as well as the median open-weight model of its size class. That is an independent lab, a standardised harness, and a comparison set the vendor did not choose. It is the most credible single number on this page.

The Number That Cuts Against It

Here is the result almost no launch coverage mentions, and it is the reason this article exists: GLM 5.3 Flash is slow.

At 50.2 output tokens per second against a 67 t/s median for its size class, it is meaningfully below average on throughput. The name is actively misleading. In most vendors' naming conventions "Flash" signals a latency tier; here it signals a price tier. The model is cheap and near-frontier-smart, and it produces tokens at a leisurely pace.

The partial consolation is time to first token: 1.47s beats the 2.14s class median, so the model starts responding quickly and then streams slowly. For a chat interface, that combination is more tolerable than the reverse. For a batch pipeline where wall-clock throughput determines your job duration, it is a real cost that does not show up in the per-token price.

This is exactly the kind of thing you should verify on your own workload rather than take on trust. Run a GLM 5.3 Flash session on glm5.app with a prompt that produces a long response and time it — five minutes of your own measurement beats any table, including this one.

How to Read These Numbers for a Decision

Differentiator: benchmark scores answer "how good," but your decision needs "good enough for what, at what cost." Here is the framework we use.

If you are migrating from GLM-5.2: the DeepSWE (46.2 → 63.4) and AutomationBench (26.2 → 48.8) deltas at roughly one-tenth the price make this close to a free upgrade. Test on your hardest task, then move.

If you are considering it against a flagship: the gap on Terminal-Bench 2.1 is 0.7 points, which sounds like nothing. Benchmarks compress at the top, though — a 0.7-point gap on aggregate scores usually reflects a larger gap on the hardest decile of tasks. If your failures are expensive to detect, keep the flagship for the hard path and route the easy 80% to Flash.

If you are weighing it against DeepSeek V4 Flash: the head-to-head on all five axes is in GLM 5.3 Flash vs DeepSeek V4 Flash.

If your workload is multimodal: OfficeQA Pro is the number to trust, and native multimodality is a structural advantage over retrofitted vision. Test it directly.

If throughput is your bottleneck: the benchmark numbers are irrelevant to you. 50.2 tok/s is the number that matters, and it is below median. Look elsewhere or budget for parallelism.

What Nobody Has Benchmarked Yet

Two gaps in the public record, stated plainly so you do not go looking for numbers that do not exist:

Long-context degradation is unmeasured. The model advertises 1,048,576 tokens and uses a hybrid KDA-linear plus NoPE-sparse-MLA attention stack with IndexPool compression to make that affordable. Z.ai reports roughly 3x less attention compute and a 4.4x smaller KV cache than GLM-5.3. What no published benchmark shows is how retrieval accuracy holds up at 800K tokens versus 80K. Compression trades precision for cost by definition; the size of that trade at this model's extremes is currently unknown.

Quantized quality is unmeasured. Community GGUF quantizations exist and are the only realistic path to self-hosting on fewer than several nodes. Every benchmark on this page was run on the full-precision FP8 release. A 4-bit quant of a model whose long-context behaviour already relies on compression is a different model, and nobody has published the delta.

If either of those matters to your deployment, treat it as an open question and test it, rather than assuming the launch numbers transfer. When you are ready to put a real task in front of it, open a GLM 5.3 Flash chat or wire the glm-5.3-flash model ID into your evaluation harness.

Frequently Asked Questions

Is GLM 5.3 Flash better than Claude Opus 4.8? On Z.ai's chosen benchmarks it is very close — 84.3 vs 85.0 on Terminal-Bench 2.1, 29.0 vs 29.5 on Z.ai's own Code Bench — but "very close on aggregate benchmarks" is not "equivalent on hard tasks." Both of those comparisons come from Z.ai. Treat it as a strong value alternative, not a flagship replacement.

What is the Artificial Analysis Intelligence Index score? 57, versus a median of roughly 27 for open-weight models of comparable size. This is the most credible number available because it is independently measured with a comparison set the vendor did not pick.

Is GLM 5.3 Flash fast? No — not in the throughput sense. 50.2 output tokens per second is below the 67 t/s median for its class. It starts fast (1.47s TTFT, better than the 2.14s median) and then streams slowly. "Flash" refers to price, not speed.

How much better is it than GLM 5.2? Substantially, on Z.ai's own within-family evaluations: DeepSWE v1.1 from 46.2 to 63.4, AutomationBench from 26.2 to 48.8 — at roughly one-tenth the price. Within-family comparisons are the most trustworthy part of the vendor table.

Are these benchmarks reproducible? Partly. Terminal-Bench 2.1, DeepSWE, HLE and AutomationBench are third-party benchmarks you can run against the open MIT-licensed weights yourself. Z.ai Code Bench v1.0 is Z.ai's own and cannot be independently verified.

Sources

Benchmark figures verified August 27, 2026. First-party results are labelled as such throughout; where a benchmark was authored by the model's vendor, that is noted in the table commentary.

Start Using GLM 5 Today

Try GLM 5 free — reasoning, coding, agents, and image generation in one platform.