AI Models

GLM-5.3 Matches Kimi K3 at 60 on Independent Intelligence Index

Z.ai's GLM-5.3 scored 60 on the Artificial Analysis Intelligence Index, matching Moonshot AI's Kimi K3 and trailing Anthropic's Claude Opus 5 at 63. The independent evaluation, published August 18, 2026, confirms GLM-5.3's post-training gains. GLM-5.3 offers the lowest cost per task among top models, and Z.ai plans to release its weights two weeks after launch, potentially reshaping the open-weights landscape.

Neura News

Neura News

Neura Market Editorial

August 18, 20266 min read
GLM-5.3 Matches Kimi K3 at 60 on Independent Intelligence Index

Z.ai's GLM-5.3 has scored 60 on the Artificial Analysis Intelligence Index, matching Moonshot AI's Kimi K3 and trailing Anthropic's Claude Opus 5, which leads at 63. The independent evaluator published the score on August 18, 2026, marking the first independent read on the model since its release on August 14, 2026. The three-point gap between GLM-5.3 and the leader is the distance a fast post-training cycle did not close.

The score places GLM-5.3 eighth in its comparison class of 181 models, where the median score sits at 35. That puts the model well above the pack, though still short of the top. The evaluation used maximum reasoning effort, a setting recommended for coding tasks. Artificial Analysis runs its own evaluation suite, so the number carries weight because the evaluator does not rely on vendor-reported results.

The Headline: Parity with Kimi K3

The parity with Kimi K3 is the headline comparison. Both models now sit at 60 on the index, but they arrived there through different routes. Kimi K3, released July 16, 2026, is the top-scoring open-weights model in Artificial Analysis rankings. Moonshot opened Kimi K3's weights under a revenue-tiered license in July 2026. GLM-5.3 shares the score but not the license, as it remains proprietary for now.

Z.ai has committed to releasing GLM-5.3's weights two weeks after the August 14, 2026 launch. That release will come after safety evaluation and hardening. If the weights drop as promised, GLM-5.3 would sit alongside Kimi K3 as an open-weights option at the 60 mark. The weight release would also put GLM-5.3 at roughly half the per-token price of Kimi K3, a gap that could reshape how developers choose between the two.

Post-Training Gains, No New Pretraining

GLM-5.3 uses the same base model as GLM-5.2, with all capability gains coming from post-training. Z.ai's release post describes a month of scaling reinforcement learning on long-horizon task environments. The training stack was built for GLM-5.2, so the gains came without starting from scratch.

The vendor-reported improvements are notable. Terminal-Bench 3.0 improved from 4.6 to 28.3, and DeepSWE v1.1 improved from 46.2 to 66.9. Z.ai documented results across coding, cybersecurity, and agentic benchmarks, comparing against Kimi K3, Claude Opus 4.8, Claude Fable 5, and GPT-5.6 Sol. Unite.AI covered the launch and the cybersecurity results separately.

These numbers are vendor-reported, which is why the Artificial Analysis figure matters more. The evaluator runs its own suite, so the 60 score is an independent check on Z.ai's claims. The Intelligence Index v4.1.1 aggregates nine evaluations, including agentic real-world work tasks, agentic tool use, terminal coding, scientific reasoning and knowledge, graduate-level science questions, physics reasoning, knowledge reliability and hallucination, and long-context reasoning.

Cost Per Task Favors GLM-5.3

Price is where the two 60-scorers diverge sharply. GLM-5.3 costs $1.40 per million input tokens and $4.40 per million output tokens on Z.ai's API. Kimi K3 costs $3.00 per million input tokens and $15.00 per million output tokens on Moonshot's API. That is a wide gap for models that score identically.

Per Intelligence Index task, GLM-5.3 comes in at $0.68, Kimi K3 at $0.84, and Claude Opus 5 at $2.34. GLM-5.3 reaches its score at the lowest cost per task of the three. The total evaluation cost for GLM-5.3 was $1,238.50 on Z.ai's API, a figure that reflects the model's pricing and its verbosity.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

GLM-5.3 is the most verbose of the group, generating 170 million output tokens across the evaluation suite. The median output tokens in its class is 72 million. That verbosity drives up raw token spend, but the per-task cost still comes out lowest because the underlying prices are so much lower than the competition. For a developer running a typical batch of 4000 tasks, the difference between $0.68 and $0.84 per task adds up to $640 in savings, and against Claude Opus 5's $2.34, the gap widens to $6,640. The price advantage holds even when accounting for the extra tokens GLM-5.3 produces.

Model Specs and Availability

Artificial Analysis records 753 billion parameters for GLM-5.3. The model is available through Z.ai's API and has rolled out to all GLM Coding Plan subscribers. The model requires thinking to be enabled, with three effort levels available. Applications calling with thinking disabled will fail until they are migrated, a warning Z.ai has posted for developers.

Claude Opus 5, released July 24, 2026, leads the index at 63. The three-point gap between GLM-5.3 and the leader shows how close the field has become. A single post-training cycle closed much of the distance but did not erase it. The comparison class median of 35 puts the top models in a different tier entirely, but the race at the top is tight.

What the Weight Release Would Change

Z.ai's commitment to releasing weights after safety evaluation and hardening is the next milestone. If the release happens, GLM-5.3 would join Kimi K3 as an open-weights option at the 60 mark. The per-token price difference would make GLM-5.3 the cheaper choice at roughly half of Kimi K3's rates.

The weight release would also change the competitive picture. Kimi K3 has held the open-weights crown since July 16, 2026. GLM-5.3 matching its score while undercutting its price would give developers a real choice. The revenue-tiered license on Kimi K3 adds another variable, though the details of Z.ai's license terms remain unannounced.

For now, GLM-5.3 is proprietary. The score of 60 is the first independent confirmation that Z.ai's post-training approach works. The model matches the top open-weights option and trails the overall leader by three points. The cost advantage makes it the cheapest path to that score, and the promised weight release could make it the most accessible one.

The next few weeks will show whether Z.ai follows through on the weight release. If it does, the open-weights tier at the 60 mark will have two players, and the price gap between them will be hard to ignore. If it does not, Kimi K3 remains the only open option at that level, and GLM-5.3 stays a proprietary alternative with a cost edge.

The evaluation cost of $1,238.50 for GLM-5.3 reflects the model's verbosity and its pricing. The 170 million output tokens it generated across the suite are more than double the class median of 72 million. That verbosity is a trade-off, but at $0.68 per task, it is a cheap one. For a developer weighing a $250 budget against a $700 one, the per-task difference of $0.16 between GLM-5.3 and Kimi K3 means roughly 80% more tasks completed for the same spend.

The field at the top is crowded. Claude Opus 5 leads at 63, with GLM-5.3 and Kimi K3 tied at 60. The median of 35 across 181 models shows how far the leaders have pulled ahead. The race now turns on price, openness, and the next round of post-training gains.

Related on Neura Market

More from Neura News

Industry

Unitree Robotics Soars 629% on Shanghai Debut, Founder Wang Xingxing Now Worth $16 Billion

Unitree Robotics shares surged as much as 629% on their Shanghai debut, making founder Wang Xingxing a billionaire with a net worth of $16 billion. The IPO raised 6.1 billion yuan, with retail demand oversubscribed by over 5,500 times. Despite geopolitical risks from a U.S. ban on Chinese humanoids, the company's revenue grew over 300% last year, and it continues to innovate with products like the high-speed 'Superman' robot and the GD01 transformable robot.

Aug 19·4 min read
Industry

OpenAI Expands Ads Pilot to 31 European Markets as Revenue Climbs

OpenAI is expanding its advertising pilot to 31 new European markets, including Germany, France, Spain, and Italy, as ad revenue grows over 25% since August 2025. With ChatGPT now reaching one billion weekly users and 20% showing commercial intent, the company is positioning itself as a major player in digital advertising. The move, announced on August 19, 2025, signals a strategic push beyond subscriptions.

Aug 19·3 min read