List rate
$0.37
Per 1M input
Standard prompt token rate.
GLM 5.3 FlashX (glm-5.3-flashx / glm5.3-flashx) is Zhipu AI's 320B MoE model (18B active) with accelerated throughput, 1M context, and OpenAI API support. Chat free online or integrate with your developer stack.
Zhipu AI's high-speed inference variant of GLM 5.3 Flash for low-latency coding and agent workflows.
Serving-stack acceleration cuts multi-turn response latency.
Sparse MoE activating 8 of 288 experts per token: 320B depth with 18B speed.
Native multimodal comprehension across 1,048,576 tokens with 128K output.
Architecture details, context limits, and open-weight checkpoints for glm-5.3-flashx from Zhipu AI.
Architecture specs
Sparse MoE across 45 layers.
8 of 288 experts per token.
Repository-scale context.
Up to 128K completion.
Vision in, text out.
Low-latency decode tier.
Deployment & download
Commercial use allowed.
Mirrored on glm 5.3 flashx huggingface and glm-5.3-flashx modelscope.
Official glm-5.3-flashx fp8 checkpoint and glm 5.3 flashx gguf.
Run glm 5.3 flashx ollama commands.
Standard endpoints.
Function calling & JSON.
Specifications confirmed from model cards and API telemetry.
Pricing / 04
Transparent token pricing and glm5.app plans. Try free online or scale production workloads.
List rate
$0.37
Standard prompt token rate.
List rate
$1.25
High-throughput completion rate.
Prompt caching
$0.09
Up to 75% savings on repeated prefixes.
glm5.app plan
Free Trial
Daily credits plus trial for glm5.3-flashx on sub plan.
Benchmarks / 05
Empirical evaluation on code synthesis, terminal agents, and token delivery.
Throughput benchmark
Faster multi-turn decode throughput.
SWE coding benchmark
Resolves GitHub engineering issues.
Agent evaluation
Matches frontier agentic execution.
Live coding benchmark
Outscores DeepSeek V4.1 Flash in coding.
Capabilities / 07
Key developer workflows.
Sub-second completions and debugging.
Ingest whole codebases without chunking.
Function calling with parseable JSON schemas.
Inspect screenshots alongside stack traces.
Deploy locally with Ollama, vLLM, GGUF, or FP8.
Standard streaming with prompt caching.
Comparison / 06
Compare architecture (diff in glm-5.3 flash and flashx / glm 5.3 flashx und flash unterschiede) and pricing (Glm 5.3 flashx vs glm 5.3 f).
High-speed flash
Low-latency streaming
8 of 288 active experts
128K max output
Cached at $0.09
Vision in, text out
Cost-efficiency
Budget batch processing
Same base weights
128K max output
Lowest class pricing
Vision in, text out
Competitor
Budget baseline
Lightweight MoE
Smaller 128K window
Commodity pricing
No vision input
Flagship prime
Frontier intelligence
Frontier scale
128K max output
Premium tier
Multimodal input
Official provider rates as of October 2026.
Use Cases / 08
Where high throughput delivers maximum developer value.
Fast completions without typing lag.
Responsive multi-turn tool and terminal execution.
Real-time streaming with 1M context.
High-speed parsing of PDFs into JSON.
Synthesize React components from UI images.
Quantized inference behind private VPCs.
Three quick routes to get started.
Chat free on glm5.app with the model preset.
Send standard OpenAI chat completions.
Download weights or run 'ollama run glm-5.3-flashx'.
Integration specimen
Call the model using standard OpenAI client libraries or cURL. Supports streaming and tool declarations.
API referencecurl https://glm5.app/api/v1/chat/completions \
-H "Authorization: Bearer $GLM5_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5.3-flashx",
"messages": [{
"role": "user",
"content": "Optimize this Python async handler for low-latency streaming completions."
}],
"stream": true
}'Answers covering release details, pricing, Ollama deployment, and multimodal capabilities.
High-speed 320B MoE model (18B active) with 1M context and accelerated throughput.
Both share 320B-A18B MoE weights and 1M context. FlashX runs on an accelerated stack for higher throughput ($0.37/$1.25 per 1M tokens).
320B total and 18B active per token across 45 layers, routing 8 of 288 experts.
1,048,576 tokens (1M) with up to 131,072 tokens (128K) completion limit.
No. It accepts multimodal vision inputs, but outputs text. Use CogView on glm5.app for image generation.
Weights are on Hugging Face (glm 5.3 flashx huggingface) and ModelScope (glm-5.3-flashx modelscope). Run locally via 'ollama run glm-5.3-flashx' (glm 5.3 flashx ollama), GGUF (glm 5.3 flashx gguf), or FP8 (glm-5.3-flashx fp8).
List rates are $0.37/1M input and $1.25/1M output ($0.09 cached). Subscriptions include discounted credits for glm5.3-flashx on sub plan.
It provides 1M context and vision inputs; DeepSeek V4.1 Flash is text-only with 128K context.
Flash is budget MoE ($0.15/$0.50), FlashX is high-speed MoE ($0.37/$1.25), GLM 5.3 is the 745B flagship, and Prime is the accelerated flagship.
Yes. Users receive free daily credits. Following the launch (智谱宣布 glm-5.3-flashx 正式上线并开启双周体验活动), users can also apply for the two-week trial.
Start here
High-speed coding, agent loops, and multimodal analysis free on glm5.app.