Chinese lab Zhipu (Z.ai) announced GLM-5.3 on 14 August, describing it as the strongest open-source model for programming and claiming a 50% improvement over GLM-5.2 on its internal hands-on coding evaluation.

The number that stands out

Terminal Bench 3.0 goes from 4.6 to 28.3 — a sixfold move on the benchmark that measures whether a model can actually operate a shell over a long task. Agentic terminal work has been the widest gap between Chinese open-weight models and Western frontier systems, and Zhipu is claiming most of it closed in one release.

The rest of the card

DeepSWE v1.1 rises to 66.9 from 46.2; Agents' Last Exam to 28.5 from 23.8; AutomationBench to 48.2 from 26.2. On Z.ai Code Bench at the High setting the model reports 31.4% accuracy while spending roughly 50,000 tokens per task — a cost figure most vendors omit, and one worth reading next to the accuracy.

Read the same chart again

Zhipu's own comparison puts GLM-5.3 behind GPT-5.6 Sol and Fable 5 on all six benchmarks it published — Terminal Bench 28.3 against 34.6 and 33.7. On DeepSWE it also sits just below Kimi K3 at 67.5. The claim is "strongest open-source", and on one of its own six benchmarks even that is contested.

No new pretraining run

Zhipu says GLM-5.3 uses the same base model as GLM-5.2 and that the gains come entirely from scaling post-training. That is the same claim xAI made for Grok 4.6 two days earlier, and it points at the same conclusion: the expensive part of a model's life is increasingly what happens after pretraining ends.

What is not verifiable yet

Every figure here is self-reported, and the weights are not published. Zhipu says they will follow in about two weeks, after safety evaluation and hardening, with API access shortly after; the company's Hugging Face organisation still lists GLM-5.2 as its latest release. Until the weights land, "strongest open-source coding model" is a claim about a model nobody outside Zhipu can run.