Zhipu AI releases GLM 5.3 FlashX high-speed tier

GLM 5.3 flashx

GLM 5.3 FlashX (glm-5.3-flashx / glm5.3-flashx) is Zhipu AI's 320B MoE model (18B active) with accelerated throughput, 1M context, and OpenAI API support. Chat free online or integrate with your developer stack.

Interactive GLM 5.3 FlashX Playground

What is GLM 5.3 FlashX

Zhipu AI's high-speed inference variant of GLM 5.3 Flash for low-latency coding and agent workflows.

01

High-speed inference stack

Serving-stack acceleration cuts multi-turn response latency.

02

320B MoE with 18B active

Sparse MoE activating 8 of 288 experts per token: 320B depth with 18B speed.

03

1M context & vision perception

Native multimodal comprehension across 1,048,576 tokens with 128K output.

GLM 5.3 FlashX parameters (GLM-5.3-Flash parameters) & open weights

Architecture details, context limits, and open-weight checkpoints for glm-5.3-flashx from Zhipu AI.

Architecture specs

Architecture parameters & limits

Total parameters
320B

Sparse MoE across 45 layers.

Active parameters
18B

8 of 288 experts per token.

Context window
1,048,576 tokens

Repository-scale context.

Maximum output
131,072 tokens

Up to 128K completion.

Input modalities
Text, Image, Video

Vision in, text out.

Serving tier
Accelerated Flash

Low-latency decode tier.

Deployment & download

Glm 5.3 flashx download: HuggingFace, ModelScope, GGUF & FP8

Weights license
MIT / Open Weights

Commercial use allowed.

Hugging Face / ModelScope
THUDM / zai-org

Mirrored on glm 5.3 flashx huggingface and glm-5.3-flashx modelscope.

Quantization formats
FP8, GGUF, AWQ

Official glm-5.3-flashx fp8 checkpoint and glm 5.3 flashx gguf.

Local deployment
Ollama, vLLM, SGLang

Run glm 5.3 flashx ollama commands.

API compatibility
OpenAI-compatible

Standard endpoints.

Tool calling
Tools & JSON

Function calling & JSON.

Specifications confirmed from model cards and API telemetry.

Pricing / 04

GLM 5.3 FlashX pricing: API rates & subscriptions

Transparent token pricing and glm5.app plans. Try free online or scale production workloads.

01

List rate

$0.37

Per 1M input

Standard prompt token rate.

02

List rate

$1.25

Per 1M output

High-throughput completion rate.

03

Prompt caching

$0.09

Per 1M cached

Up to 75% savings on repeated prefixes.

04

glm5.app plan

Free Trial

Daily chat & 2-week trial

Daily credits plus trial for glm5.3-flashx on sub plan.

Standard provider rates. Subscribers receive metered discounts.

Benchmarks / 05

GLM 5.3 FlashX benchmarks and coding performance

Empirical evaluation on code synthesis, terminal agents, and token delivery.

01

Throughput benchmark

Decode throughput

FlashX95.0
Standard Flash50.0

Faster multi-turn decode throughput.

02

SWE coding benchmark

DeepSWE v1.1 (%)

FlashX65.0
GLM 5.246.0

Resolves GitHub engineering issues.

03

Agent evaluation

Terminal-Bench 2.1 (%)

FlashX85.0
Opus 4.885.0

Matches frontier agentic execution.

04

Live coding benchmark

LiveCodeBench pass@1 (%)

FlashX55.0
DeepSeek V4 Flash51.0

Outscores DeepSeek V4.1 Flash in coding.

Scores from official evaluation leaderboards.

Capabilities / 07

Core capabilities and development workflows of GLM 5.3 FlashX

Key developer workflows.

01

Real-time interactive code completion

Sub-second completions and debugging.

02

1M-token repository-scale context analysis

Ingest whole codebases without chunking.

03

Precision tool calling

Function calling with parseable JSON schemas.

04

Multimodal visual inputs

Inspect screenshots alongside stack traces.

05

Self-hosting: Ollama, GGUF & FP8

Deploy locally with Ollama, vLLM, GGUF, or FP8.

06

OpenAI-compatible API

Standard streaming with prompt caching.

Comparison / 06

GLM 5.3 FlashX vs GLM 5.3 Flash vs DeepSeek V4.1 Flash vs Prime

Compare architecture (diff in glm-5.3 flash and flashx / glm 5.3 flashx und flash unterschiede) and pricing (Glm 5.3 flashx vs glm 5.3 f).

High-speed flash

GLM 5.3 FlashX

Primary strength
High-speed throughput

Low-latency streaming

Architecture
320B total / 18B active MoE

8 of 288 active experts

Context window
1M tokens (1,048,576)

128K max output

API price (Input/Output)
$0.37 / $1.25 per 1M

Cached at $0.09

Image generation (glm 5.3 flashx image generation)
Input only (no image gen)

Vision in, text out

Cost-efficiency

GLM 5.3 Flash

Primary strength
Lowest token cost

Budget batch processing

Architecture
320B MoE (18B active)

Same base weights

Context window
1M tokens (1,048,576)

128K max output

API price (Input/Output)
$0.15 / $0.50 per 1M

Lowest class pricing

Image generation
Input only

Vision in, text out

Competitor

DeepSeek V4.1 Flash

Primary strength
Budget reasoning

Budget baseline

Architecture
Latent MoE

Lightweight MoE

Context window
128K tokens

Smaller 128K window

API price (Input/Output)
$0.14 / $0.28 per 1M

Commodity pricing

Image generation
Text only

No vision input

Flagship prime

GLM 5.3 Prime

Primary strength
Deep reasoning power

Frontier intelligence

Architecture
745B MoE

Frontier scale

Context window
1M tokens (1,048,576)

128K max output

API price (Input/Output)
$1.40 / $4.40 per 1M

Premium tier

Image generation
Input only

Multimodal input

Official provider rates as of October 2026.

Use Cases / 08

Production use cases optimized for GLM 5.3 FlashX

Where high throughput delivers maximum developer value.

01

Interactive IDE coding assistants

Fast completions without typing lag.

02

Autonomous agent execution & tool chaining

Responsive multi-turn tool and terminal execution.

03

Live customer support

Real-time streaming with 1M context.

04

Large-scale document extraction

High-speed parsing of PDFs into JSON.

05

UI-to-code generation

Synthesize React components from UI images.

06

Private on-premise deployments

Quantized inference behind private VPCs.

Quickstart / 09

How to use GLM 5.3 FlashX in 3 easy steps

Three quick routes to get started.

01
01

1. Try free online

Chat free on glm5.app with the model preset.

02
02

2. Call the OpenAI-compatible API

Send standard OpenAI chat completions.

03
03

3. Run locally via Ollama or GGUF

Download weights or run 'ollama run glm-5.3-flashx'.

Integration specimen

Start with the GLM 5.3 FlashX API

Call the model using standard OpenAI client libraries or cURL. Supports streaming and tool declarations.

API reference
Request previewcurl
curl https://glm5.app/api/v1/chat/completions \
  -H "Authorization: Bearer $GLM5_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3-flashx",
    "messages": [{
      "role": "user",
      "content": "Optimize this Python async handler for low-latency streaming completions."
    }],
    "stream": true
  }'

GLM 5.3 FlashX FAQ: Answers to common developer questions

Answers covering release details, pricing, Ollama deployment, and multimodal capabilities.

What is GLM 5.3 FlashX (glm-5.3-flashx / glm5.3-flashx)?

High-speed 320B MoE model (18B active) with 1M context and accelerated throughput.


What is the diff in GLM-5.3 Flash and FlashX (glm 5.3 flashx vs glm 5.3 flash / und Flash Unterschiede)?

Both share 320B-A18B MoE weights and 1M context. FlashX runs on an accelerated stack for higher throughput ($0.37/$1.25 per 1M tokens).


How many parameters does it have (GLM 5.3 flashx parameters / GLM-5.3-Flash parameters)?

320B total and 18B active per token across 45 layers, routing 8 of 288 experts.


What is the context window of FlashX?

1,048,576 tokens (1M) with up to 131,072 tokens (128K) completion limit.


Can GLM 5.3 FlashX generate images (glm 5.3 flashx image generation)?

No. It accepts multimodal vision inputs, but outputs text. Use CogView on glm5.app for image generation.


Where can I find Glm 5.3 flashx download links and run locally via Ollama, GGUF, or Hugging Face?

Weights are on Hugging Face (glm 5.3 flashx huggingface) and ModelScope (glm-5.3-flashx modelscope). Run locally via 'ollama run glm-5.3-flashx' (glm 5.3 flashx ollama), GGUF (glm 5.3 flashx gguf), or FP8 (glm-5.3-flashx fp8).


What is the pricing on API and subscription plans (glm5.3-flashx on sub plan / glm 5.3 flashx pricing)?

List rates are $0.37/1M input and $1.25/1M output ($0.09 cached). Subscriptions include discounted credits for glm5.3-flashx on sub plan.


How does it compare to DeepSeek V4.1 Flash (glm 5.3 flashx vs deepseek v4.1 flash)?

It provides 1M context and vision inputs; DeepSeek V4.1 Flash is text-only with 128K context.


What is the difference between GLM 5.3 Flash, FlashX, and Prime (glm 5.3 flashx vs prime)?

Flash is budget MoE ($0.15/$0.50), FlashX is high-speed MoE ($0.37/$1.25), GLM 5.3 is the 745B flagship, and Prime is the accelerated flagship.


Is it free to use on glm5.app (glm 5.3 flashx free)?

Yes. Users receive free daily credits. Following the launch (智谱宣布 glm-5.3-flashx 正式上线并开启双周体验活动), users can also apply for the two-week trial.


Start here

Start building with GLM 5.3 FlashX today

High-speed coding, agent loops, and multimodal analysis free on glm5.app.