AI Models

Alibaba's Qwen3.8 Max matches Claude Opus 4.8, but Kimi K3 wins on price

Alibaba's Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, matching Claude Opus 4.8 but trailing Kimi K3's 57. However, Kimi K3 is 25% cheaper per task, and Qwen3.8 Max shows regressions in honesty and long-context retrieval, with higher hallucination rates.

Neura News

Neura News

Neura Market Editorial

August 6, 20263 min read
Alibaba's Qwen3.8 Max matches Claude Opus 4.8, but Kimi K3 wins on price

Alibaba's new Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, a 10-point jump over its predecessor's 46. That puts the model on par with Claude Opus 4.8, ahead of GLM-5.2's 51, and just one point behind Kimi K3's 57. The catch: Kimi K3 runs 25 percent cheaper.

A big leap, with a rival in sight

The jump from Qwen3.7 Max to Qwen3.8 Max is substantial. The older model scored 46 on the same index, so the new release gains a full 10 points. Alibaba now sits level with Claude Opus 4.8, a notable achievement for the Chinese technology company. But Kimi K3 remains the leader, and it undercuts Alibaba on cost.

On the GDPval-AA benchmark, which measures work-related tasks, Qwen3.8 Max leaps 468 Elo points to 1,739. Kimi K3 scores 1,685, while Claude Opus 5 leads the field at 1,852. The improvement is real, yet the efficiency gap tells a different story.

Thorough work comes at a price

Qwen3.8 Max needs 64 steps per task on GDPval-AA. Kimi K3 needs just 14. That difference drives up costs in a hurry. Input tokens grew 15x because the test resends full conversation history at each step. The model works more thoroughly, but it runs slower and costs more.

Alibaba's price-to-performance ratio takes a hit as a result. The company cut token prices, but the per-task cost still rose. Input token price dropped from $2.50 to $2.00 per million tokens. Output token price fell from $7.50 to $6.00 per million tokens. Cache hit price went from $0.50 to $0.25 per million tokens. Those cuts sound good, but they do not offset the extra work.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

A single task in the Intelligence Index now costs $1.14 for Qwen3.8 Max. Qwen3.7 Max costs $0.53 per task. Kimi K3 costs $0.86 per task and scores one point higher. GLM-5.2 costs $0.57 per task. The math is clear: Alibaba's new model is the most expensive of the group, and it does not lead the index.

Honesty and long-context regressions

Performance gains come with trade-offs. AA-LCR, which tests whether a model can correctly pull together information from very long texts, dropped 2 points compared to the previous version. AA-Omniscience, which measures whether a model answers knowledge questions correctly or honestly admits it doesn't know, fell 10 points.

The accuracy rate stays around 31 percent. The hallucination rate jumped from 23 to 40 percent. Qwen3.8 Max guesses far more often instead of saying it doesn't know. That is a troubling shift for a model that otherwise improved so much on raw intelligence.

What the numbers mean for buyers

The benchmark results, published by Artificial Analysis on Aug 6, 2026, paint a mixed picture. Qwen3.8 Max catches Claude Opus 4.8, but Kimi K3 still scores higher for 25 percent less. Alibaba's price cuts on tokens are real, but the model's extra steps and larger inputs erase those savings.

For teams that value raw capability, Qwen3.8 Max is a strong option. For those watching costs, Kimi K3 looks better. And for anyone relying on honest answers or long-context retrieval, the regressions are a warning sign. The model is faster to guess, and that guess comes with a price tag of its own.

Related on Neura Market

More from Neura News

AI Models

42 Mathematicians Urge Royal Society to Warn Government and Media About AI Existential Risk

Forty-two mathematical fellows, including Fields Medal winners Martin Hairer, Peter Scholze, and Wendelin Werner, have signed an open letter urging the Royal Society to warn the UK government and media about existential risks from advanced AI. The letter follows recent breakthroughs in which leading models solved open research problems, including a Millennium Problem. None of the signatories are affiliated with AI companies. The group warns that AI labs' estimates of existential risk above ten percent must not be dismissed as hype, and that by the time the situation becomes obvious to the public, it may be too late to act.

Sep 18·2 min read
Developer

Steve Yegge Shuts Down Gas Town After Failing to Build Anything Else With It

Steve Yegge shut down Gas Town, his ultra-vibed coding agent orchestrator, after admitting he never built anything else with it despite heavy subscription spend. Databricks reported a 60% coding spend increase after rolling out GPT-6 Astra to 3,500 engineers, OpenAI published a misalignment disclosure framework with six case reports, and Xiaomi ran MiMo-V2.6 RL training in public with live telemetry.

Sep 18·21 min read