
Dev Log /
GLM-5.3 Just Shipped, and It Is Not a Small Update
- Models
Z.ai shipped GLM-5.3 on August 14 as what it describes as a post-training-only upgrade: the same base model as GLM-5.2, trained over more environments, more diverse tasks, and more compute. The benchmark jumps do not read like a routine post-training pass. Z.ai reports a 50% improvement over GLM-5.2 on its own internal Code Bench, plus open-source state-of-the-art results on Terminal Bench 3.0 and Agents’ Last Exam.
The same base, trained much harder
Post-training-only means no new pre-trained foundation — the upgrade lives entirely in what the model was taught afterward. The environments Z.ai built for this round were designed to resemble real units of expert work rather than short coding exercises: some tasks represent several days of work for an experienced engineer, complete with compute clusters, storage systems, internal documentation, and a working codebase, each verified end to end before it ever reaches the reward signal.
The results of that environment scaling, from Z.ai’s own comparison table:
Terminal Bench 3.0 jumps from 4.6% to 28.3% — on the benchmark that most directly tests long-running terminal work, the previous score was effectively a failure and the new one is competitive. DeepSWE rises from 46.2% to 66.9%, AutomationBench nearly doubles from 26.2% to 48.2%, and Humanity’s Last Exam with tools climbs from 54.7% to 62.5%.
More score for fewer tokens
The upgrade also shows up as efficiency. At high effort, GLM-5.3 hits 31.4% accuracy on Z.ai’s Code Bench using roughly 50,000 output tokens per task — beating Claude Opus 4.8’s best result of 29.5% at maximum effort, which needs about 120,000 tokens to get there. A higher score on less than half the tokens is the kind of improvement that compounds: every agent loop, every CI run, every batch job pays it back per call.
The unexpected result: security
The strangest gain was not on the coding benchmarks at all. Z.ai added vulnerability-discovery data to the post-training mix expecting modest improvements in spotting individual bugs; instead the model started reasoning across entire exploitation chains, forming coherent plans that carry from initial discovery through full exploitation. CyberGym climbs from 77.2% to 84.5% — the highest score of any model in Z.ai’s comparison set — and ExploitBench more than doubles, from 24.4% to 54.4.
Set loose on real open-source codebases with security-team collaborators, GLM-5.3 helped surface 2,436 vulnerabilities across 269 projects, 1,097 of them critical or high severity. Some dated back to 1981 and had sat undiscovered for an average of 26.6 years. Z.ai is publishing the findings through a public Security Disclosure Ledger as they clear responsible-disclosure timelines.
The gap to the closed frontier
Against GPT-5.6 Sol, the distance that was wide on GLM-5.2 is now close, and in places it has closed entirely. GLM-5.3 leads on AutomationBench, 48.2 to 45.8, and on CyberGym, 84.5 to 83.6, and sits within a point of Sol on Agents’ Last Exam, 28.5 to 28.6. Sol still leads on Terminal Bench 3.0 and DeepSWE, but the gap shrank on every benchmark tested.
This is the closest an open-weights model has come to the closed frontier on agentic coding — and the first time one has led on live vulnerability-finding rather than a leaderboard score. Weights are expected publicly about two weeks after launch, once Z.ai’s safety-hardening pass on the release is complete.