GLM 5.3 Cybersecurity: CyberGym SOTA, 2,436 Real Vulnerabilities & the Disclosure Ledger

GLM 5.3 Cybersecurity: CyberGym SOTA, 2,436 Real Vulnerabilities & the Disclosure Ledger

GLM 5.3's emergent cyber capability — CyberGym 84.5 (best public result), 2,436 real-world vulnerabilities found across 269 projects, and why Z.AI holds the weights for safety hardening.

GLM 5.3 Cybersecurity: CyberGym SOTA, 2,436 Real Vulnerabilities & the Disclosure Ledger

Quick answer: GLM 5.3's most surprising result isn't coding — it's security. The model scores 84.5 on CyberGym, the best public vulnerability-discovery result (ahead of Mythos 5's 83.8 and GPT-5.6 Sol's 83.6), and in real-world testing with Chinese security teams it identified 2,436 vulnerabilities across 269 projects — 1,097 of them medium-to-high severity, some hidden for decades. This emergent capability is exactly why Z.AI is holding the open weights for two weeks of safety hardening.


TL;DR

QuestionAnswer
CyberGym score84.5 — best public result on the benchmark
vs GPT-5.6 Sol83.6 — GLM 5.3 leads
Real-world findings2,436 vulnerabilities, 269 projects
Medium-to-high severity1,097 (107 critical, 990 high)
Oldest flaw foundIntroduced in 1981 (~45 years)
DisclosureZ.ai Security Disclosure Ledger (53 public, 2,383 under embargo)
Why weights are delayedSafety evaluation + hardening for exactly this capability

The "Emergent" Capability Nobody Expected

Z.AI is candid about what happened: the team added vulnerability-discovery data and environments to the post-training mix, expecting the model to get better at finding flaws. Instead, as training scaled, the capability "developed faster than we expected" — GLM 5.3 started reasoning across complete exploitation chains, not just isolated bugs.

"GLM-5.3 did not simply become better at identifying isolated flaws: it began to reason across multiple stages of exploitation, forming coherent plans for complete exploitation chains."

That's the difference between a model that flags a suspicious line and one that can map a full attack path — and it's the reason the open-weight release is gated.

CyberGym: Best Public Result

BenchmarkGLM 5.3GLM 5.2Kimi K3Mythos 5GPT-5.6 Sol
CyberGym84.577.280.083.883.6
ExploitBench54.424.432.278.076.5
ExploitGym (2h / 6h)105 / 13029 / 3936 / 70181 / 247216 / 293

CyberGym starts from white-box source code and tests whether the model can identify and validate vulnerabilities by triggering faults. GLM 5.3's 84.5 leads the entire comparison — ahead of both Mythos 5 and GPT-5.6 Sol.

The pattern Z.AI highlights: the further up the exploitation chain a benchmark sits, the larger GLM 5.3's gain over GLM 5.2 — ExploitBench more than doubled (24.4 → 54.4), ExploitGym tripled (29 → 105 tasks in 2h). Capability grew fastest exactly where the dual-use risk concentrates — and where GLM still trails the closed frontier (Mythos 5: 78.0 ExploitBench).

Real-World Results: 2,436 Vulnerabilities

Benchmarks transfer — Z.AI has been working with Chinese security teams against real codebases since GLM 5.2:

  • 2,436 vulnerabilities found across 269 projects
  • 1,097 medium-to-high severity (107 critical, 990 high, 1,286 medium, 53 low)
  • Scope: system kernels, operating systems, browser engines, open-source infrastructure, web applications, network protocols
  • 45 years of impact — the oldest flaw was introduced in 1981; on average a vulnerability lived 26.6 years before discovery

The findings flow through a Z.ai Security Disclosure Ledger — a public record tracking each issue from discovery through disclosure: affected project, severity, CVE where available, and how long the flaw had existed. As of launch: 53 publicly disclosed, 2,383 under embargo.

Defensive Positioning — and the Policy Question

Z.AI frames the capability as defensive: finding and patching vulnerabilities before attackers exploit them, with a structured disclosure pipeline. The numbers support that framing — 1,097 medium-to-high findings handed to maintainers through the ledger.

But the honesty is refreshing: the same capability, applied offensively, is a weapon. That's the tension in the launch — and why the weights wait. Z.AI's stated plan is two weeks of safety evaluation and hardening before the HuggingFace release.

If you deploy GLM 5.3 for security work:

  1. Expect strong vuln discovery — CyberGym 84.5 SOTA means real triage throughput in code review pipelines.
  2. Budget for the disclosure workflow — a model that finds 1,097 medium-to-high issues needs a CVE process, not a to-do list.
  3. Write your own guardrails — Z.AI's two-week hold covers its evaluation, not your deployment policy.

FAQ

Is GLM 5.3 good at finding vulnerabilities? Yes — the best public CyberGym result (84.5), and 2,436 real-world vulnerabilities found across 269 projects in testing with security teams.

Why is Z.AI delaying the open weights? Safety evaluation and hardening, specifically because GLM 5.3's cyber capability — vulnerability discovery plus exploitation-chain reasoning — emerged faster than expected.

Is GLM 5.3 better than GPT-5.6 Sol at security? On CyberGym, yes (84.5 vs 83.6). On ExploitBench and ExploitGym, GPT-5.6 Sol still leads (76.5 vs 54.4; 216/293 vs 105/130).

What is the Z.ai Security Disclosure Ledger? A public record of vulnerabilities GLM found in real projects, tracking severity, CVE, and disclosure status (53 public, 2,383 under embargo at launch).

Is the cyber capability defensive or offensive? Z.AI positions it defensively (find and patch before attackers exploit), but acknowledges the dual-use risk — the reason for the gated weight release.


Sources

Last updated: August 14, 2026

Start Using GLM 5 Today

Try GLM 5 free — reasoning, coding, agents, and image generation in one platform.