GLM-5.3 and Flash: Benchmarks, Architecture, and Limits
GLM-5.3 and Flash: Benchmarks, Architecture, and Limits
Meta description: GLM-5.3 claims open-weight coding SOTA and found 2,436 real vulnerabilities. Flash runs at 1/10 the cost. The benchmarks, the architecture, and the gaps.
Before anyone knew its name, GLM-5.3-Flash spent six days as the most-used model on OpenRouter. Listed under the anonymous alias "ox-alpha," it processed 23.2 trillion tokens — 2.3 times the runner-up — while every request ran on Chinese-made AI accelerators. When Zhipu AI pulled the curtain on August 25, 2026, two models emerged, not one.
GLM-5.3 is a 744B coding and cybersecurity flagship. GLM-5.3-Flash is a 320B efficiency model with just 18B active parameters, open-sourced under the MIT license.
Both claim open-weight state of the art on coding benchmarks. Both were built by scaling post-training — more reinforcement learning, harder tasks, longer horizons — rather than growing the base model. Flash delivers roughly 90% of GLM-5.3's coding performance for one-tenth the API price, and adds native vision that GLM-5.3 lacks.
GLM-5.3-Flash weights are already on ModelScope and Hugging Face. GLM-5.3 weights followed on August 28 after a two-week safety review prompted by the model's unexpectedly strong vulnerability-finding abilities — Zhipu wanted to evaluate those capabilities before making the weights self-hostable.
Post-training scaling: the shared method
GLM-5.3 and GLM-5.3-Flash share the post-training stack Zhipu built during the GLM-5.2 cycle: IndexShare for long-context processing, SAO with compression for long-horizon reinforcement learning, and slime for large-scale asynchronous RL training. All three run on continuously accumulated long-horizon task environments — executable worlds where a model must plan, act, and iterate over extended sessions to earn a reward.
The fork is architectural. GLM-5.3 takes GLM-5.2's 744B Mixture-of-Experts base and pushes post-training further — more environments, more tasks, more compute. No new parameters. GLM-5.3-Flash starts from a brand-new 320B base designed for inference efficiency, then passes through the same pipeline.
GLM-5.3: pushing post-training to the limit
Training on production-scale engineering tasks
GLM-5.3's training environments go well beyond coding puzzles. Each task represents a complete engineering unit — some equivalent to several days of work for a senior engineer. A machine learning infrastructure task gives the model the same workspace a human gets: compute cluster, storage, internal docs, codebase, and experiment logs. The job is to diagnose a training-stack bottleneck, implement a fix, run experiments, and deliver measurable speedup.
Research agents start the synthesis chain by collecting task patterns from real engineering work and converting them into runnable long-horizon environments with multi-step dependencies and hidden state. Judge agents verify each task is solvable. Verifiers are built without seeing reference solutions, and solver trajectories close reward shortcuts. The binary signals that come out are clean enough to train on directly.
Top-p masking, full-vocabulary OPD, and R3-style training-inference consistency went into the slime RL framework — reducing logprob divergence to 1e-7 levels, more than 99.99% lower than before. Joint scheduling and load-balancing for long-horizon workloads lifted end-to-end RL training throughput by over 2.3x.
Coding: open-weight records, but not absolute SOTA
Terminal-Bench 3.0 jumped from 4.6 to 28.3. DeepSWE v1.1 went from 46.2 to 66.9. Agents' Last Exam rose from 23.8 to 28.5. All three are open-weight records.
On Zhipu's internal Z.ai Code Bench, GLM-5.3 at "Max" effort scores 34.5% with roughly 75K output tokens per task. GLM-5.2 managed 23.4% at 96K tokens — so the newer model is both more accurate and more token-efficient. At "High" effort, GLM-5.3 hits 31.4% at about 50K tokens, compared to Claude Opus 4.8's 29.5% at 120K.
Against closed-source models, the picture shifts. GLM-5.3's Terminal-Bench score of 28.3 trails GPT-5.6 Sol at 34.6 and Claude Fable 5 at 33.7. On DeepSWE, its 66.9 is behind GPT-5.6 Sol's 72.7 and Fable 5's 69.7.
"Strongest open-weight coding model" is a defensible claim. "Strongest coding model" is not.
Cybersecurity: strong at finding bugs, weaker at exploiting them
Vulnerability discovery data and environments were part of GLM-5.3's post-training. As training scaled, the model went beyond spotting isolated bugs — it started reasoning across exploit chains, forming multi-stage attack plans. Zhipu says the capability growth outpaced expectations.
Three benchmarks capture different stages of the security pipeline.
CyberGym (white-box source-code vulnerability identification): 84.5%, the top score on this benchmark — ahead of Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%. All three fall within one percentage point.
ExploitBench (deep exploitation reasoning): 54.4%, more than double GLM-5.2's 24.4%. But Mythos 5 scores 78.0% and GPT-5.6 Sol 76.5% on the same test.
ExploitGym (time-normalized exploitation tasks): 105 completed in 2 hours, 130 in 6 hours, versus GLM-5.2's 29 and 39. Mythos 5 reaches 181 and 247.
The pattern: GLM-5.3's cybersecurity strength concentrates on the upstream end — finding vulnerabilities in source code, where it genuinely leads among all models. The further downstream toward writing working exploits, the wider the gap to closed-source frontier models. For defensive security research, that upstream strength is exactly where the value sits.
2,436 real vulnerabilities, mostly still under embargo
Working with multiple Chinese security teams, Zhipu tested GLM-5.3 against real codebases. After expert review and deduplication: 2,436 vulnerabilities across 269 open-source projects, 1,097 rated medium or high severity. The affected projects span system kernels, operating systems, browser engines, web applications, and network protocols.
Many are ancient. The oldest traces back to code introduced around 1981 — roughly 45 years, with an average age across all findings of 26.6 years.
At the time of the announcement, 53 had been publicly disclosed with CVEs assigned. The remaining 2,383 sit under embargo while maintainers patch. That embargo is the direct reason Zhipu held back GLM-5.3's weights — a model that finds decades-old bugs this effectively needed a safety review before anyone could self-host it.
GLM-5.3-Flash: new architecture, one-tenth the cost
Hybrid linear-sparse attention
320B total parameters, but only 18B active per token across 45 layers — down from GLM-4.5's 32B active across 92 layers. The gap between total and active is where the efficiency lives.
Linear attention handles local dependencies through state modeling. Sparse attention retrieves global context through a lightweight indexer. At million-token context lengths, the indexer's latency and memory costs become the bottleneck, so Zhipu introduced IndexPool — weighted pooling that compresses four indexer key vectors into one. Manifold-constrained hyperconnection (mHC) further improves scaling.
The numbers: 3.0x lower attention compute and 4.4x smaller KV cache compared to GLM-5.3. Among all models Zhipu benchmarked — including DeepSeek-V4-Flash and Kimi-K3 — GLM-5.3-Flash has the lowest attention compute.
Pre-trained on 30 trillion multimodal tokens, the base model scores 37.6 on LiveCodeBench. GLM-4.5-Base manages 28.1, DeepSeek-V4-Flash-Base 29.9.
Native multimodal vision
GLM-5.3-Flash is the first GLM-5 model with built-in vision — not routed through a separate adapter, but integrated directly into the inference loop. The model decides when to look and uses what it sees to guide its next action.
Zhipu's training pipeline builds on self-visual-judgment: trajectories where the model interacts with an environment, inspects its own output, and iterates until the result passes its own check. For frontend coding, they added RL with environment feedback, verifying against real user flows to sharpen GUI judgment. The result handles documents, spreadsheets, dashboards, and slide decks, and evaluates rendered output for both correctness and visual quality.
All of it runs on Chinese AI chips
Every ox-alpha request — 23.2 trillion tokens on OpenRouter, roughly 42 trillion on OpenCode — ran on domestically produced Chinese AI accelerators. The inference stack sits on SGLang and includes tensor parallelism for linear attention and the LM head, ReplaySSM, W8A8 quantization, mixed INT8/FP8/BF16 cache quantization, and Layer Split.
At the cluster level, an EPD (Encode-Prefill-Decode) architecture separates multimodal encoding, prompt prefilling, and token-by-token decoding into independently scheduled worker pools. End-to-end performance improved 3x over the baseline on identical hardware, reaching per-token costs comparable to mainstream NVIDIA GPUs.
The ox-alpha stealth test was more than a marketing play. It was a production-scale proof that Chinese chips can serve a frontier model under real traffic without users noticing a difference.
Choosing between them
GLM-5.3-Flash scores 57 on the Artificial Analysis Intelligence Index (v4.1.1) at $0.09 per task — about $0.045 at Zhipu's promotional pricing. That intelligence level previously cost roughly 10x more.
The cost drop is already reshaping how software gets built: AI-native app builders like HappySeeds let users ship production apps through natural conversation — a workflow that only pencils out when the underlying model is capable enough for real code generation and cheap enough to call thousands of times per session.
On hard coding benchmarks, GLM-5.3 keeps a modest lead: DeepSWE +3.5, Terminal-Bench 2.1 +3.9. On agent tasks, Flash pulls ahead by +5.4 on Toolathlon. Vision is Flash territory alone.
Pick GLM-5.3 for maximum coding depth and security research where every percentage point on CyberGym matters. Pick GLM-5.3-Flash for production workloads, multimodal tasks, and any deployment where cost per token outweighs the last few points on a coding leaderboard.
- GLM-5.3-Flash on ModelScope (MIT license) | Technical blog
Comments
Post a Comment