Z.ai announced the release of GLM-5.3, a model that reuses the same 743 billion parameter base as GLM-5.2 and concentrates all improvements in post-training. The company said the weights will be published in about two weeks, after the security assessment and system hardening are completed.

Coding performance

On Terminal-Bench 3.0, GLM-5.3 went from 4.6 to 28.3 points compared with GLM-5.2. On DeepSWE v1.1, the score went from 46.2 to 66.9. On Agents' Last Exam (CLI), the advance was from 23.8 to 28.5. On GDPval-AA v2, which covers 44 occupations, the model scored 1,769 points.

GLM 5.3 Performance Evaluation
GLM 5.3 Performance Evaluation

On Z.ai's internal Code Bench benchmark, the company reported a 50% improvement over GLM-5.2, with 31.4% accuracy on tasks with about 50,000 output tokens. The test compares the model with Claude Opus 4.8 (29.5% with 120,000 tokens) and Claude Fable 5 (39.5% at maximum effort). The company says the private benchmark reduces the risk of contamination.

On public suites, GLM-5.3 trails GPT-5.6 Sol and Fable 5 in some of the harder coding evaluations. All results are self-reported by the developers, with configurations documented in the announcement.

Cybersecurity results

The company classified the cybersecurity advance as unplanned. The company added vulnerability discovery data hoping to improve reasoning about isolated bugs, but the capability continued to expand with training, forming coherent plans in full exploitation chains.

Cyber Capability
Cyber Capability

On CyberGym, which tests discovery and validation from source code, GLM-5.3 rose from 77.2% to 84.5%, surpassing Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%). On ExploitBench, which requires reasoning about root cause and functional exploitation, the model went from 24.4% to 54.4%, still below Mythos 5, with 78.0%.

On ExploitGym, GLM-5.3 completed 105 tasks in two hours and 130 in six hours, versus 29 and 39 for GLM-5.2. Mythos 5 completed 181 and 247 tasks in the same periods.

Availability

GLM-5.3 is available through the Z.ai API, the GLM Coding Plan, and ZCode. The company recommends that startups and midsize engineering teams adopt the model immediately, while companies with data residency rules or vendor review requirements should wait for the release of the weights.

More from Radar