GLM-5.3 and Flash: Benchmarks, Architecture, and Limits
GLM-5.3 and Flash: Benchmarks, Architecture, and Limits Meta description: GLM-5.3 claims open-weight coding SOTA and found 2,436 real vulnerabilities. Flash runs at 1/10 the cost. The benchmarks, the architecture, and the gaps. Before anyone knew its name, GLM-5.3-Flash spent six days as the most-used model on OpenRouter. Listed under the anonymous alias "ox-alpha," it processed 23.2 trillion tokens — 2.3 times the runner-up — while every request ran on Chinese-made AI accelerators. When Zhipu AI pulled the curtain on August 25, 2026, two models emerged, not one. GLM-5.3 is a 744B coding and cybersecurity flagship. GLM-5.3-Flash is a 320B efficiency model with just 18B active parameters, open-sourced under the MIT license . Both claim open-weight state of the art on coding benchmarks. Both were built by scaling post-training — more reinforcement learning, harder tasks, longer horizons — rather than growing the base model. Flash delivers roughly 90% of GLM-5.3's coding pe...