Z.ai GLM-5.3 Boosts Coding and Cyber Skills, Flags Reported Cursor Flaw
Z.ai's GLM-5.3 boosts coding and cyber benchmarks through post-training, but a reported Cursor flaw and delayed open weights sharpen AI safety concerns.
Summary
On August 14, 2026, Chinese startup Z.ai, formerly Zhipu AI, released GLM-5.3 through GLM Coding Plan and ZCode for macOS, Windows and Linux; API access and open weights follow safety hardening, with weights targeted in two weeks and pricing undisclosed. Scaling post-training instead of replacing GLM-5.2's roughly 743-billion-parameter base, it scored 28.3, 66.9, 48.2 and 28.5 on Terminal-Bench 3.0, DeepSWE v1.1, AutomationBench and Agents' Last Exam CLI, up from 4.6, 46.2, 26.2 and 23.8, while trailing GPT-5.6 Sol and Claude Fable 5 on the first two. Private Z.ai Code Bench showed Max at 34.5% using 75,000 output tokens, versus GLM-5.2's 23.4% at 96,000; High hit 31.4% at 50,000, versus Claude Opus 4.8's 29.5% at 120,000.
Cyber gains prompted controls and Reuters-reported “trusted access.” GLM-5.3 scored 84.5% on CyberGym, ahead of GLM-5.2's 77.2%, GPT-5.6 Sol's 83.6% and Mythos 5's 83.8%; ExploitBench rose from 24.4% to 54.4%, below Sol's 76.5% and Mythos 5's 78%. ExploitGym completions rose from 29 to 105 under two hours and 39 to 130 under six, behind Fable 5's 181 and 247 and Sol's 216 and 293. Z.ai says Chinese security teams produced 2,436 reviewed, deduplicated findings across 269 projects, including 1,097 critical or high severity, 53 public and 2,383 embargoed. Developer advocate Lou said GLM-5.3 found a potentially serious Cursor vulnerability; Cursor, recently acquired by SpaceX, had not confirmed it.
Migration requires thinking enabled, with low, high or max effort; max is default, and thinking.type "disabled" requests fail. Promotional plans cost $12.60 monthly for Lite with 10,000 weekly credits, $56 for Pro, $117.60 for Max, and $88 or $188 per Team seat; off-peak calls use 50% points. The leap shows post-training can extend coding agents without new pretraining, while faster-than-expected exploit capability complicates open release.
Positives
- Terminal-Bench 3.0 rose from 4.6 to 28.3, while DeepSWE v1.1 climbed from 46.2 to 66.9.
- Max reasoning achieved 34.5% on Z.ai Code Bench using 75,000 output tokens, improving on GLM-5.2's 23.4% with 96,000.
- CyberGym reached 84.5%, exceeding Z.ai's reported results for GPT-5.6 Sol at 83.6% and Mythos 5 at 83.8%.
- 2,436 reviewed and deduplicated vulnerability findings across 269 projects demonstrate practical security-research output.
- Open weights are targeted roughly two weeks after launch, once Z.ai completes safety evaluation and hardening.
Risks & concerns
- Cursor had not confirmed developer advocate Lou's claim that GLM-5.3 found a potentially serious vulnerability in its software.
- ExploitBench capability more than doubled to 54.4%, raising offensive-use concerns despite trailing GPT-5.6 Sol and Mythos 5.
- 2,383 of Z.ai's 2,436 vulnerability findings remained under embargo, including findings among 1,097 classified as critical or high severity.
- API access, open weights and general API pricing remain unavailable while Z.ai conducts staged safety work.
- Applications sending thinking.type "disabled" will fail until developers enable thinking and select low, high or max reasoning effort.