AI Models

Microsoft's New Copilot Model Trails DeepSeek on Benchmarks and Price

Microsoft's new MAI Code 1.1 Flash model for GitHub Copilot underperforms DeepSeek's open-weight alternative on Terminal Bench 2.1 while costing more per token. Independent benchmarks show DeepSeek-V4-Flash-0731 scoring 82.7% versus Microsoft's 62.9%, despite Microsoft's claims of improved efficiency. The gap raises questions about Microsoft's open AI positioning as it pushes proprietary models.

Neura News

Neura News

Neura Market Editorial

August 12, 20263 min read
Microsoft's New Copilot Model Trails DeepSeek on Benchmarks and Price

Microsoft released MAI Code 1.1 Flash on Aug 12, 2026, a proprietary code model for GitHub Copilot that the company says writes better code than its predecessor. Independent benchmarks tell a different story. The new model underperforms DeepSeek's open-weight alternative on key tests while costing more per token, a gap that raises fresh questions about Microsoft's public embrace of open AI.

What Microsoft Claims

Microsoft says MAI Code 1.1 Flash is 25% more token-efficient than its June predecessor and costs a quarter as much. The company also reports that developers accepted 4% more of its output compared with the earlier version, and that "code survival rose 4% and return visits increased 9%." Training involved "hundreds of thousands of reinforcement-learning environments in GitHub Copilot."

The official announcement touts these vague improvement metrics but skips direct comparisons with rivals. Microsoft's model card buries the benchmark results, according to The Decoder's analysis, leaving the headline numbers to carry the message.

The Benchmark Gap

On SWE-bench Verified, MAI Code 1.1 Flash scores 72.6%, edging past its predecessor's 71.6% and beating Anthropic's Claude Haiku 4.5 at 69.8% and OpenAI's GPT-5.4 mini at 69.2%. DeepSeek-V4-Flash-0731's SWE-bench Verified score is not published.

The Terminal Bench 2.1 results are more revealing. MAI Code 1.1 Flash scores 62.9%, up from 51.7% for MAI Code 1 Flash. Claude Haiku 4.5 trails at 49.4%, and GPT-5.4 mini reaches 60.7%. DeepSeek-V4-Flash-0731 crushes them all with an 82.7% score.

The Price Problem

Pricing widens the gap. DeepSeek-V4-Flash charges $0.14 per input token, $0.0028 per input token with cache, and $0.28 per output token. MAI Code 1.1 Flash costs $0.20 per input token, $0.02 per input token with cache, and $1.20 per output token. Claude Haiku 4.5 is pricier still at $1.00 per input token, $0.10 per input token with cache, and $5.00 per output token.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Cost per token doesn't tell the whole story without factoring in usage efficiency, but the gap in DeepSeek's favor is likely significant either way. MAI Code 1.1 Flash looks like a budget model on paper, yet it trails the more capable DeepSeek on both price and performance.

Open Talk, Closed Models

Microsoft has been positioning itself as an open AI champion. Its model strategy tells a different story. MAI Code 1.1 Flash is proprietary and likely won't get an open-weights release. The company recently shook up Copilot, swapping out OpenAI and Anthropic models for its own cheaper MAI alternatives to cut costs.

The trade-off was worse performance for better margins. Matthias Bastian, writing for The Decoder, argues that Microsoft is sinking resources into a weaker, pricier in-house model instead of tapping more capable, freely available alternatives like DeepSeek-V4-Flash. Microsoft's open AI talk doesn't match its model strategy.

What Comes Next

Customers in the Microsoft ecosystem can still pick from different models depending on the app and use case. That flexibility may not last. Microsoft will almost certainly make its own models the default eventually, according to the analysis. That would lock up a massive share of the market, since most users never actively choose a specific AI model anyway.

The pattern is clear. Microsoft's proprietary push prioritizes margins over capability, and its communication hides unfavorable comparisons. For developers who care about output quality and cost, the open-weight alternative from DeepSeek remains the stronger choice.

Related on Neura Market

More from Neura News

Research

Modular Pretraining: A New Approach to Containing Dangerous AI Knowledge

Researchers at Anthropic and AE Studio have introduced Gradient Routed Auxiliary Modules (GRAM), a method that isolates dangerous knowledge in large language models into switchable modules during training. This approach allows operators to control access to sensitive content, potentially reducing risks of misuse. Preliminary experiments show promise across models up to 5B parameters, but the method has not yet been applied to production-scale systems.

Aug 17·12 min read
Industry

AI and Data Centers Overtake Israel and Racism as Top US Campaign Issues

A Washington Post analysis reveals AI and data centers have become leading issues in the 2026 US midterm elections, mentioned in nearly 40% of House, Senate, and governor races. The topic now outranks Israel, manufacturing, and racism, with Democrats focusing on regulation and child safety while Republicans emphasize national security and competition with China. Public skepticism about AI and job losses is driving the political focus.

Aug 17·4 min read