GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash — the same 320B-parameter, 18B-active Mixture-of-Experts architecture with a 1M-token context window, tuned for inference speeds of up to 200 tokens per second. It launched on OpenRouter under the model ID z-ai/glm-5.3-flashx on September 18, 2026, and lists at $0.37 per 1M input tokens and $1.25 per 1M output tokens — roughly 2.5x the price of the standard Flash tier.
If you have been researching this model, you have probably hit the same confusion everyone does. The GLM-5.3 family now ships three tier names — GLM 5.3, GLM 5.3 Flash, and GLM 5.3 FlashX (plus a GLM-5.3-Prime for the flagship) — and "FlashX" appears in search results under a dozen spellings: glm-5.3-flashx, glm 5.3 flashx, even the typo glm 5.3flahx. On top of the name collision, the headline speed claim ("up to 200 tokens/s") does not match what infrastructure monitors actually measure as a median, and the price premium over Flash is not explained anywhere in plain arithmetic.
This article fixes both problems. It is built from the OpenRouter model page and its public endpoints data (queried October 7, 2026), the Hugging Face model card for the shared architecture, and Artificial Analysis's independent speed measurements for the Flash tier. Where a number is a vendor claim rather than an independent measurement, we label it as such. We have not run our own benchmark suite against GLM-5.3-FlashX, and we say so instead of dressing up vendor charts as testing.
GLM-5.3-FlashX at a Glance
| Field | Value |
|---|---|
| Model ID | z-ai/glm-5.3-flashx (OpenRouter) |
| Vendor | Z.ai (Zhipu AI) |
| Released | September 18, 2026 |
| Relationship to Flash | High-speed variant of GLM-5.3-Flash |
| Architecture | Same as Flash: 320B total / 18B active MoE, hybrid sparse + linear attention |
| Modality | Text, image, and video input →text output |
| Context window | 1,048,576 tokens (1M) |
| Max output | 131,072 tokens |
| Claimed speed | Up to 200 tokens/s (vendor/OpenRouter description) |
| Measured speed (P50) | 50 tok/s output, 2.27s round-trip latency (OpenRouter monitoring) |
| Price on OpenRouter | $0.37 / 1M input, $1.25 / 1M output, $0.09 / 1M cache read |
| Tool calling | Yes (tools and tool_choice) |
| Structured output | response_format JSON, without JSON-schema enforcement |
What FlashX Actually Is (and Isn't)
The one-sentence version: FlashX is GLM 5.3 Flash with an inference-speed upgrade, not a new model. The OpenRouter description is explicit that it is "the high-speed variant of Z.ai's GLM-5.3-Flash," built on the same hybrid sparse and linear attention architecture — the 320B-total, 18B-active design Z.ai shipped on August 26, 2026. Your prompts see the same model weights and the same native multimodal pathway (text, image, and video in; text out). What changes is the serving configuration: Z.ai accelerated the inference stack, and prices it as a separate model tier.
The pattern is not unique to Flash. Z.ai applies the same split across the family: GLM-5.3-Prime is the high-speed variant of the GLM 5.3 flagship, described on OpenRouter as inheriting full capabilities while delivering roughly 1.5x the output throughput through inference acceleration. FlashX is the same idea applied to the Flash tier. Once you understand that mapping, the whole GLM-5.3 lineup snaps into a 2x2 grid: two model bases (GLM 5.3 flagship and GLM 5.3 Flash) times two serving tiers (standard and high-speed).
What FlashX is not: a bigger context window, a multimodal upgrade, or a new reasoning mode. Context stays at 1,048,576 tokens with a 131,072-token completion ceiling — the same window Flash advertises. If you need more output headroom, that is a separate question from speed, and neither Flash nor FlashX currently lists more than 131K completion tokens on OpenRouter.
The Price Premium, in Plain Arithmetic
FlashX costs more than Flash for the same weights. Here is the exact premium using list rates:
| Direction | GLM 5.3 Flash (list) | GLM-5.3-FlashX (OpenRouter) | Premium |
|---|---|---|---|
| Input / 1M | $0.15 | $0.37 | ~2.5x |
| Cached input / 1M | $0.03 | $0.09 | ~3x |
| Output / 1M | $0.50 | $1.25 | ~2.5x |
Two caveats matter for budgeting. First, Flash's $0.15/$0.50 is the undiscounted list price; a temporary 50% launch discount ($0.075/$0.25) has been honoured by some providers, which makes the effective FlashX premium larger whenever the discount is live on your endpoint — we break the discount trap down in our GLM 5.3 Flash pricing guide. Second, OpenRouter currently lists one provider for FlashX, so there is no price-shopping spread the way there is on Flash, where ten providers serve the same weights at different rates. If that changes, the premium arithmetic changes with it.
A worked example makes the premium concrete. Take a representative agentic coding task — 120K input tokens, 15K output tokens:
- On Flash at list rates: 0.12M × $0.15 + 0.015M × $0.50 ≈$0.026 per task
- On FlashX: 0.12M × $0.37 + 0.015M × $1.25 ≈$0.063 per task
That is about 2.4x the spend for the identical task. Across 10,000 daily agent runs, the gap is roughly $370 versus $630 per day. Whether that premium is justified depends entirely on one question: does the speed difference actually shorten your wall-clock time? Which brings us to the part most coverage gets wrong.
Speed: the "200 Tokens/s" Claim vs. What's Measured
The headline claim — inference speeds of up to 200 tokens/s — is a vendor-supplied description that OpenRouter reproduces on the model page. It is labelled "up to" for a reason, and OpenRouter's own public performance monitoring tells you what typical traffic actually sees: a P50 (median) throughput of 50 tok/s and a 2.27s round-trip latency on the Z.ai provider, measured across real production traffic as of October 7, 2026.
Medians understate peaks, so 200 tok/s bursts and 50 tok/s medians are not necessarily contradictory — serving engines often hit high throughput on long, uncontested generation streams and lower throughput on short, tool-heavy agent turns where the model stops to think. But the honest reading for a buyer is this:
- For streaming chat UIs, budget around the median, not the peak. If perceived speed is the product, prototype with the median and confirm the peak on your workload shape.
- For agent loops, time-to-first-token often matters more than tok/s. Flash's independently measured time to first token was already strong (1.47s per Artificial Analysis), and if FlashX's acceleration mainly targets decode throughput, agent workflows dominated by tool calls gain less than streaming chat does.
- For batch jobs, tok/s × cost is the whole story. Batch throughput gains shrink wall-clock time, but at 2.4x the token cost, the math only works if speed is actually your bottleneck.
This is the same discipline we recommend for Flash itself, where Artificial Analysis measures 50.2 output tokens per second — below the 67 tok/s class median — despite the "Flash" name. "Flash" means cheap, and "FlashX" means faster, but neither name tells you what your workload will measure. Run your own task before committing: you can try GLM 5.3 Flash free in glm5.app chat with screenshots and real prompts in a five-minute test.
When FlashX Is Worth It — and When It Isn't
FlashX makes sense when:
- You run latency-sensitive streaming products where Flash's ~50 tok/s median measurably hurts retention or completion rates, and you have confirmed on your own traffic that the faster variant delivers.
- Your workload is long-form generation — long code diffs, long documents — where decode throughput dominates wall-clock time and the 2.5x premium buys a higher throughput.
- Your per-task value is high enough that compute time costs more than the token premium. A coding agent that finishes in half the time and bills $0.063 instead of $0.026 is cheaper per completed task if it frees a $0.10+ of engineer-minute equivalent.
Standard Flash makes sense when:
- You run high-volume, cost-dominated workloads — classification, extraction, bulk summarization — where a 2.4x spend increase is never recovered.
- Your agent loops spend most of their wall-clock time on tool calls, retrieval, or waiting, not on decoding tokens. Accelerating decode does nothing for those.
- You want provider price-shopping leverage. Flash is served by many providers at a 2x spread (see the ten-provider table in GLM 5.3 Flash pricing); FlashX currently is not.
Neither fits when the task needs flagship depth. Both Flash tiers are efficiency models. For the hardest sliver of tasks where a wrong answer is expensive to detect, the GLM 5.3 flagship tier still earns its premium — our GLM 5.3 vs GLM 5.3 Flash comparison walks through that decision.
Features You Keep (and One You Partially Lose)
Because FlashX serves the same base model as Flash, the agentic feature set carries over:
- Function calling via
toolsandtool_choice— supported, per the OpenRouter model page. - Structured output via
response_formatfor parseable JSON — supported, though without JSON-schema enforcement, so your application should still validate responses. - Native multimodal input — screenshots, diagrams, and screen recordings alongside text, inherited from Flash's 30T-token multimodal pre-training corpus.
- Million-token context — with the cache-read rate at $0.09 per 1M tokens. Note that Flash's cache rate is $0.03, so cache-heavy, long-context workloads pay 3x more per cached token on FlashX — a compounding cost for repository-scale agents that keep a stable prefix resident.
The partial loss is pricing flexibility: with one provider listed, there is currently no provider.order routing game to play the way there is on Flash. If your pipeline pins providers for price or context ceiling, check availability before designing around it.
How to Start with GLM-5.3-FlashX
Three routes, in order of effort:
- Try it in the browser first. glm5.app chat runs GLM 5.3 Flash with free credits for new accounts — validate quality on your real tasks before paying any tier's premium.
- Call the API. glm5.app exposes an OpenAI-compatible endpoint; swap the model ID per tier. For the Flash tier, use
glm-5.3-flashathttps://glm5.app/api/v1— full examples in the GLM 5.3 Flash API guide. On OpenRouter, the FlashX ID isz-ai/glm-5.3-flashx. - Self-host the base weights. Flash's MIT-licensed weights are on Hugging Face at
zai-org/GLM-5.3-Flash(~328 GB in FP8, a multi-node job). FlashX's speed tier is a serving configuration; the honest self-host path is the open weights plus your own accelerated serving stack.
Start with GLM-5.3-FlashX Today
Validate before you pay the premium: open glm5.app chat free, run a streaming task that looks like your production traffic, and compare perceived speed against the GLM 5.3 Flash product page benchmarks. If the standard tier's speed is already enough, stay on Flash; if your UI needs more headroom, wire the z-ai/glm-5.3-flashx model ID into your pipeline and measure the difference on your own bill.
Frequently Asked Questions
What is GLM-5.3-FlashX?
It is the high-speed variant of Z.ai's GLM-5.3-Flash, released September 18, 2026. Same 320B-total/18B-active hybrid-attention architecture, same native multimodal input, same 1M-token context window — retuned serving for up to 200 tokens/s claimed inference speed, priced at $0.37/$1.25 per 1M input/output tokens on OpenRouter.
GLM-5.3-FlashX vs GLM-5.3-Flash — which should I use?
Flash for cost-dominated and high-volume workloads (2.5x cheaper per token, with a ten-provider price spread); FlashX when decode speed measurably improves your product's wall-clock time and you have verified the gain on your own traffic. Both keep the same capabilities: tool calling, structured JSON output, image and video input.
Is the 200 tokens/s speed real?
It is a vendor-supplied "up to" claim reproduced on the model page. OpenRouter's own public monitoring shows a P50 of 50 tok/s output with 2.27s latency on the serving provider as of October 7, 2026. Peaks and medians differ by workload shape — benchmark with your own prompts rather than trusting either number alone.
How much does GLM-5.3-FlashX cost?
$0.37 per 1M input tokens, $1.25 per 1M output tokens, and $0.09 per 1M cache read on OpenRouter, as of October 7, 2026. That is roughly 2.5x Flash's list price for the same base model.
What is the difference between FlashX and Prime?
They are the same concept applied to different tiers: GLM-5.3-Prime is the high-speed variant of the GLM 5.3 flagship (roughly 1.5x its output throughput), while FlashX is the high-speed variant of the Flash efficiency tier. Prime inherits the flagship's capabilities; FlashX inherits Flash's.
Is GLM-5.3-FlashX free?
No. The GLM 5.3 Flash base model was free during its six-day Ox Alpha preview week in August 2026, but all hosted tiers are metered now. You can still try GLM 5.3 Flash free in the browser on glm5.app with new-account credits.
Sources
Facts in this article were verified against the following primary sources on October 7, 2026:
- GLM 5.3 FlashX on OpenRouter — model description, release date, pricing ($0.37/$1.25/$0.09), context and output limits, tool-calling and structured-output support, measured P50 throughput and latency.
- zai-org/GLM-5.3-Flash on Hugging Face — the shared 320B/18B hybrid-attention architecture, native multimodal pre-training corpus, and MIT license of the base model FlashX serves.
- GLM-5.3-Flash on Artificial Analysis — independent output-speed (50.2 tok/s) and TTFT (1.47s) measurements for the base Flash tier, used as the speed context for the variant comparison.
- GLM 5.3 Flash pricing guide — our own ten-provider price table for the Flash tier, queried August 27, 2026, used for the premium arithmetic.
Prices and measured performance change frequently; confirm current rates on the OpenRouter model page before committing a production budget.




