Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks
Anthropic
OpenAI
xAI
Z.ai has released GLM-5.3, which runs on the same 743B base model as GLM-5.2. All gains come from scaled post-training, with coding benchmarks like Terminal-Bench 3.0 jumping from 4.6 to 28.3 and cybersecurity scores reaching 84.5% on CyberGym. Weights are not yet public, but the model is live via API and Coding Plan, with weights expected in about two weeks.
Z.ai released GLM-5.3, which runs on the same 743B base model as GLM-5.2, with all reported improvements coming from scaled post-training involving more task environments, more environment types, and longer training. The model shows significant gains in long-horizon coding benchmarks, with Terminal-Bench 3.0 improving from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, and Agents’ Last Exam (CLI) from 23.8 to 28.5. On the internal Z.ai Code Bench, the company reports a 50% improvement over GLM-5.2, scoring 31.4% at roughly 50,000 output tokens per task, compared to Claude Opus 4.8's 29.5% at 120,000 tokens, while Claude Fable 5 still leads at 39.5%. In cybersecurity, Z.ai describes the gains as unplanned, with CyberGym moving from 77.2% to 84.5%, surpassing Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%, and ExploitBench more than doubling from 24.4% to 54.4%, though still trailing Mythos 5 at 78.0%. The model is available via the Z.ai API, the GLM Coding Plan, and ZCode, but weights are not yet public. Z.ai says it will publish weights roughly two weeks after launch, once safety evaluation and hardening finish, and highlights that the deeper into the exploitation chain a benchmark sits, the larger the gain over GLM-5.2.
- Abbreviations
- API = Application Programming Interface — программный интерфейс приложения
- CLI = Command Line Interface — интерфейс командной строки
Source: MarkTechPost —
original
