Z.ai’s GLM-5.3 has been evaluated by Artificial Analysis at 60 on its Intelligence Index, the independent evaluator reported on August 18, 2026, placing the Chinese lab’s newest reasoning model level with Moonshot AI’s Kimi K3 and three points behind Anthropic’s Claude Opus 5, the current leader at 63.
The score covers GLM-5.3 running at its maximum reasoning effort, the setting Z.ai recommends for coding work. At 60, it sits well above the 35 median of the 181 models in its comparison class, and eighth in the class overall.
The result is the first independent read on a model that Z.ai released on August 14, 2026 with an unusual construction: GLM-5.3 uses the same base model as GLM-5.2, with every capability gain coming from post-training rather than a new pretraining run.
How GLM-5.3 Got Here
Z.ai’s release post describes a month of scaling reinforcement learning on long-horizon task environments: more environments, more diverse tasks, more compute, all on the training stack the lab built for GLM-5.2. The company reported that approach moved Terminal-Bench 3.0 from 4.6 to 28.3 and DeepSWE v1.1 from 46.2 to 66.9, and its own post documents a spread of results across coding, cybersecurity, and agentic benchmarks against Kimi K3, Claude Opus 4.8, Claude Fable 5, and GPT-5.6 Sol. Unite.AI covered the launch and its emergent cybersecurity results in detail when the model shipped last week.
The Artificial Analysis number carries different weight than those vendor-reported tables, because the evaluator runs its suite itself. Its Intelligence Index v4.1.1 aggregates nine evaluations spanning agentic real-world work tasks, agentic tool use, terminal coding, scientific reasoning and knowledge, graduate-level science questions, physics reasoning, knowledge reliability and hallucination, and long-context reasoning. GLM-5.3’s 60 is the composite of its run across that battery, at a total evaluation cost of $1,238.50 on Z.ai’s API.
Where GLM-5.3 Lands on the Leaderboard
The parity with Kimi K3 is the headline comparison. Kimi K3, released July 16, 2026, scores the same 60 on the index and remains the top-scoring open-weights model in Artificial Analysis’s rankings, a position GLM-5.3 now shares in score if not in license. Moonshot opened Kimi K3’s weights under a revenue-tiered license in July 2026; GLM-5.3 is listed as proprietary for now, with Artificial Analysis recording 753 billion parameters for the model.
Claude Opus 5, released July 24, 2026, still leads the index at 63. The three-point gap between GLM-5.3 and the leader is the distance a fast post-training cycle did not close, against a frontier that has itself moved since the spring.
Price is where the two 60-scorers diverge sharply. GLM-5.3 costs $1.40 per million input tokens and $4.40 per million output tokens on Z.ai’s API, against $3.00 and $15.00 for Kimi K3 on Moonshot’s. Per Intelligence Index task, that works out to $0.68 for GLM-5.3 versus $0.84 for Kimi K3 and $2.34 for Claude Opus 5. GLM-5.3 reaches its score at the lowest cost per task of the three, though it is also the most verbose of the group, generating 170 million output tokens across the evaluation suite against a 72 million median in its class.
What Ships Next
GLM-5.3 is available through Z.ai’s API and has been rolled out to all GLM Coding Plan subscribers. The model requires thinking to be enabled, with three effort levels, and Z.ai warns that applications still calling it with thinking disabled will fail until migrated.
The weights are the remaining piece. Z.ai has committed to releasing them two weeks after the August 14, 2026 launch, once safety evaluation and hardening are complete, which would put GLM-5.3 alongside Kimi K3 as an open-weights option at the 60 mark on the index, at roughly half the per-token price.
Credit: Source link



























