GLM 5.3 Pricing: Coding Plan Points, API Cost & Off-Peak Discounts

GLM 5.3 Pricing: Coding Plan Points, API Cost & Off-Peak Discounts

GLM 5.3 pricing explained — points-based GLM Coding Plan, 50% off-peak rate, API per-token cost, and the ZCode 1.5x quota boost ending August 31.

GLM 5.3 Pricing: Coding Plan Points, API Cost & Off-Peak Discounts

Quick answer: GLM 5.3 doesn't have a separate price tag — it's included in the GLM Coding Plan, which now uses a points-based quota (input, cached input, and output tokens metered separately). Calls made outside peak hours (14:00–18:00 UTC+8, Mon–Fri) consume 50% of standard points. API users pay per token for model id glm-5.3, and ZCode adds a limited-time 1.5× quota boost through August 31.


TL;DR

AccessPricing modelKey lever
GLM Coding PlanPoints-based quota50% points off-peak (evenings + weekends)
ZCodePlan quota + cache savings1.5× boost through Aug 31; 98%+ cache hit rate
API (glm-5.3)Per-token billingUse low/high reasoning effort to cut cost
Open weightsFree (self-host)Weights in two weeks — no API cost at all

GLM 5.3 Is Included in the GLM Coding Plan

Z.AI rolled GLM 5.3 out to all GLM Coding Plan subscribers on launch day (August 14, 2026). There's no separate "5.3 add-on" — if you're on the Coding Plan, you're already running it.

The plan itself changed to a points-based quota system. The key details:

  • Points are calculated separately for input, cached input, and output tokens.
  • Off-peak calls cost 50% of standard points. Peak hours are 14:00–18:00 (UTC+8), Monday through Friday; everything else — early morning, evenings, and all weekend — runs at the half-price rate.
  • If your app sends long system prompts repeatedly, the cached input metering is your biggest discount lever: cached context bills at the lower cached-input rate.

The Off-Peak Discount, Explained

This is the single most useful pricing fact in the new plan:

  • Peak: 14:00–18:00 UTC+8, Mon–Fri → full points.
  • Off-peak: every other hour, plus all weekends → 50% of standard points.

In practice: schedule batch jobs, long agent runs, and heavy refactors for evenings or weekends and you can roughly double your effective quota. For teams operating across time zones, routing non-urgent work to off-peak hours is the cheapest optimization available.

ZCode: Cache Savings + 1.5× Quota Boost (Through Aug 31)

If you use ZCode as your GLM 5.3 agent:

  • 98%+ cache hit rate — repeated context bills at the lower cached rate, which Z.AI estimates gives you ~30% more effective tokens.
  • 1.5× limited-time quota boost — through August 31, stack it with cache savings for up to 180% of your standard quota.

The boost is time-boxed, so if you're evaluating GLM 5.3 for a big migration, the window to stress-test it cheaply closes at the end of August.

API Pricing: Per Token, Effort-Controlled

The Z.AI API bills per token for glm-5.3. Since GLM 5.3 now requires thinking to be enabled (thinking.type: "enabled"), your main cost control is the effort level:

reasoning_effortBest forCost profile
lowClassification, extraction, light formattingLowest — minimal thinking tokens
highOrdinary coding, Q&A, draftingBalanced
maxComplex multi-file coding, long-horizon agentsHighest per request (but fewer retries)

Z.AI's own data shows GLM 5.3 is more token-efficient than GLM 5.2 — at Max effort it hits 34.5% task completion at ~75K output tokens vs 5.2's 23.4% at ~96K. Fewer tokens per completed task means a lower effective cost per result, not just a lower sticker price.

The Cheapest Option: Self-Host in Two Weeks

GLM 5.3 weights will be public on HuggingFace two weeks after launch (safety evaluation and hardening permitting). If your workloads are predictable and you have GPU capacity, self-hosting eliminates per-token cost entirely — the same playbook teams already use with GLM 5.2 weights.

How to Keep GLM 5.3 Costs Down — Checklist

  1. Run off-peak — 50% points outside 14:00–18:00 UTC+8 weekdays.
  2. Cache aggressively — stable system prompts and context prefixes hit cached-input rates.
  3. Right-size effortlow for simple tasks, max only where it pays.
  4. Use ZCode before Aug 31 — 1.5× boost stacks with cache savings.
  5. Plan for self-hosting — weights land in two weeks; benchmark your GPU needs now.

FAQ

Is GLM 5.3 free? GLM Coding Plan subscribers get it as part of their plan. Off-peak usage costs 50% of standard points. No separate free tier has been announced for 5.3.

How does the points-based quota work? Input, cached input, and output tokens each consume points at their own rate. Cached input is the cheapest tier.

What are the peak hours for GLM 5.3? 14:00–18:00 (UTC+8), Monday through Friday. All other hours — including weekends — get the 50% off-peak rate.

Does GLM 5.3 cost more than GLM 5.2? Same plan/API structure — no premium for the 5.3 upgrade. It's also more token-efficient per completed task.

Can I avoid API costs entirely? Yes — self-host once weights release in two weeks.

When does the ZCode 1.5× boost end? August 31, 2026.


Sources

Last updated: August 14, 2026

Start Using GLM 5 Today

Try GLM 5 free — reasoning, coding, agents, and image generation in one platform.

GLM 5.3 Pricing: Coding Plan Points, API Cost & Off-Peak Discounts - GLM 5