Alibaba's new Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, a 10-point jump over its predecessor's 46. That puts the model on par with Claude Opus 4.8, ahead of GLM-5.2's 51, and just one point behind Kimi K3's 57. The catch: Kimi K3 runs 25 percent cheaper.
A big leap, with a rival in sight
The jump from Qwen3.7 Max to Qwen3.8 Max is substantial. The older model scored 46 on the same index, so the new release gains a full 10 points. Alibaba now sits level with Claude Opus 4.8, a notable achievement for the Chinese technology company. But Kimi K3 remains the leader, and it undercuts Alibaba on cost.
On the GDPval-AA benchmark, which measures work-related tasks, Qwen3.8 Max leaps 468 Elo points to 1,739. Kimi K3 scores 1,685, while Claude Opus 5 leads the field at 1,852. The improvement is real, yet the efficiency gap tells a different story.
Thorough work comes at a price
Qwen3.8 Max needs 64 steps per task on GDPval-AA. Kimi K3 needs just 14. That difference drives up costs in a hurry. Input tokens grew 15x because the test resends full conversation history at each step. The model works more thoroughly, but it runs slower and costs more.
Alibaba's price-to-performance ratio takes a hit as a result. The company cut token prices, but the per-task cost still rose. Input token price dropped from $2.50 to $2.00 per million tokens. Output token price fell from $7.50 to $6.00 per million tokens. Cache hit price went from $0.50 to $0.25 per million tokens. Those cuts sound good, but they do not offset the extra work.
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
A single task in the Intelligence Index now costs $1.14 for Qwen3.8 Max. Qwen3.7 Max costs $0.53 per task. Kimi K3 costs $0.86 per task and scores one point higher. GLM-5.2 costs $0.57 per task. The math is clear: Alibaba's new model is the most expensive of the group, and it does not lead the index.
Honesty and long-context regressions
Performance gains come with trade-offs. AA-LCR, which tests whether a model can correctly pull together information from very long texts, dropped 2 points compared to the previous version. AA-Omniscience, which measures whether a model answers knowledge questions correctly or honestly admits it doesn't know, fell 10 points.
The accuracy rate stays around 31 percent. The hallucination rate jumped from 23 to 40 percent. Qwen3.8 Max guesses far more often instead of saying it doesn't know. That is a troubling shift for a model that otherwise improved so much on raw intelligence.
What the numbers mean for buyers
The benchmark results, published by Artificial Analysis on Aug 6, 2026, paint a mixed picture. Qwen3.8 Max catches Claude Opus 4.8, but Kimi K3 still scores higher for 25 percent less. Alibaba's price cuts on tokens are real, but the model's extra steps and larger inputs erase those savings.
For teams that value raw capability, Qwen3.8 Max is a strong option. For those watching costs, Kimi K3 looks better. And for anyone relying on honest answers or long-context retrieval, the regressions are a warning sign. The model is faster to guess, and that guess comes with a price tag of its own.

