AI Models

Kimi K3 Open Model Nears GPT-5.6 Sol and Fable 5 Performance

Kimi has released K3, a multimodal open-weight model with 2.8 trillion parameters and a one million token context window. In benchmarks, it approaches Claude Fable 5 and GPT-5.6 Sol but trails them slightly. Priced at $3 per million input tokens and $15 per million output tokens, K3 signals an end to ultra-cheap Chinese AI models, though it remains cheaper than top Western competitors.

Neura News

Neura News

Neura Market Editorial

July 16, 20265 min read
Kimi K3 Open Model Nears GPT-5.6 Sol and Fable 5 Performance

Chinese AI company Kimi has released K3, a multimodal open-weight model with 2.8 trillion total parameters, delivering performance near top proprietary systems but at a significantly higher price than its predecessor. The launch signals a shift in Chinese AI pricing, moving away from the rock-bottom costs that have defined the market.

K3 uses a mixture-of-experts architecture with 896 experts, though it activates only 16 of them at a time. The model supports a context window of one million tokens and processes images and video natively. Kimi calls K3 the first open model in roughly the 3 trillion parameter range. Full model weights are expected by July 27, 2026.

The model was trained on a cluster of 10,000 NVIDIA H100 GPUs over 90 days. Kimi says the training process consumed about 20 gigawatt-hours of electricity, roughly equal to the annual power use of 2,000 U.S. homes. The company claims K3 achieved a training efficiency of 38% model flops utilization, which it says is among the highest reported for models of this scale.

Performance: Near the Top, But Not the Leader

In Kimi's own benchmarks, K3 trails Claude Fable 5 and GPT 5.6 Sol but beats all other tested systems, including Claude Opus models and Chinese rival GLM-5.2. Kimi says the results were achieved at maximum or high thinking intensity.

K3 wins two out of six programming benchmarks and three out of six general agent tests in Kimi's evaluations. Fable 5 wins both visual agent tests. Across all 35 tests, K3 took first place about seven times. K3 beats Opus 4.8, GPT 5.5, and GLM 5.2 by a wide margin in nearly every benchmark. Three different agent systems were used in the benchmarks: KimiCode, Claude Code, and Codex, meaning conditions were not identical.

Independent testing lab Artificial Analysis gives K3 an Intelligence Index score of 57. That compares to Claude Fable 5 at 60, GPT-5.6 Sol at 59, and Claude Opus 4.8 at 56. K3 ranks fourth overall on the Artificial Analysis Intelligence Index.

On the GDPval v2 agentic task evaluation, K3 scores an Elo rating of 1,668, up from K2.6's 1,190. For context, GLM-5.2 scores 1,514, GPT-5.5 scores 1,494, Claude Opus 4.8 scores 1,600, and Claude Fable 5 scores 1,760.

K3 tops AutomationBench-AA with a 53% score. On the private long-horizon knowledge work evaluation AA-Briefcase, K3 achieves an Elo of 1,547, up 732 from K2.6. Only Claude Fable 5 scores higher on AA-Briefcase.

On the AA-Omniscience Index, which measures accuracy and hallucination, K3's accuracy rate improved from 33% to 46%. Its overall AA-Omniscience score rose to +18, up from +6. However, K3's hallucination rate increased to 51%, up from 39% for K2.6. Artificial Analysis calls K3 well-rounded, with rubric scoring and analytical quality close to Fable 5's level.

K3 also performed well on the MATH-500 benchmark, scoring 96.2%, compared to K2.6's 90.1%. On the MMLU-Pro test, K3 achieved 85.7%, up from 78.3%. On the HumanEval coding benchmark, K3 scored 92.4%, compared to 84.7% for K2.6. On the GSM8K math reasoning test, K3 scored 95.8%, up from 89.2%.

Architecture and Efficiency Gains

K3 uses a new attention architecture called Kimi Delta Attention. This enables up to 6.3x faster decoding for million-token contexts. Attention residuals boost training efficiency by roughly 25% with less than 2% extra compute overhead.

The model's mixture-of-experts design includes 896 experts, but only 16 are activated per token. This means each forward pass uses about 320 billion parameters, or roughly 11% of the total. Kimi says this design keeps inference costs manageable while maintaining high performance.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

K3's training used a novel data mixture that included 60% code, 25% text, and 15% multimodal data. The company says this ratio was chosen to optimize for the model's primary use case of software development. The training data was sourced from publicly available datasets and proprietary web crawls, totaling about 15 trillion tokens.

Kimi says K3's primary use case is long-running software development with minimal human oversight. The model analyzes large codebases, coordinates terminal tools, and stays focused across many steps. It uses a closed-loop system called "Vision in the Loop" for programming with visual feedback, which Kimi positions as a foundation for game development, UI design, and CAD.

Demos include a procedurally generated 3D open-world game built with Three.js, WebGPU, and GPU Compute, an interactive black hole visualization, a Long March 10 rocket simulation, and a Game Boy Advance emulator.

K3 also supports tool use through function calling, with native integration for Python execution, shell commands, and web browsing. The model can handle up to 10 concurrent tool calls per turn. Kimi says this makes K3 suitable for automated workflows in data analysis, DevOps, and scientific research.

Pricing: A Shift in Chinese AI Economics

K3's API pricing marks a notable increase from its predecessor. Input tokens cost $0.30 per million with a cache hit and $3.00 per million without cache. Output tokens cost $15.00 per million, including reasoning. These prices apply regardless of context length. Caching is automatic, and unmodified long prefixes are useful for agents and large codebases.

For comparison, K2.6 priced input tokens at $0.16 per million with a cache hit, $0.95 without cache, and output at $4.00 per million. K3 is much pricier than its predecessor but comparable to Western mid-range models like Claude Sonnet 5, which charges $3 per million input tokens and $15 output.

Claude Fable 5 charges $1.00 input with cache hit, $10.00 input without hit, and $50.00 output. GPT 5.6 Sol charges $0.50 input with cache hit, $5.00 input without hit, and $30.00 output.

On the Intelligence Index, K3's per-task cost is $0.94. GPT-5.6 Sol costs $1.04 per task, Opus 4.8 costs $1.80, GLM-5.2 costs $0.32, and DeepSeek V4 Pro costs just $0.04. K3 will likely still cost more per task than K2.6 in most cases despite using fewer tokens.

K3 used roughly 132 million output tokens for all nine evaluations, down from about 166 million for K2.6, a 21% reduction. K3 scored 13 points higher than K2.6 while using fewer tokens.

The pricing signals that Chinese providers are no longer offering frontier models at rock-bottom prices. K3 is available on Kimi.com, mobile apps for iOS, Android, and HarmonyOS, Kimi Work desktop version 3.1.0 and above, and Kimi Code. On OpenRouter, the model identifier is "moonshotai/kimi-k3." Open weights are expected by the end of July 2026.

Kimi offers a separate business version with member management and a personal/business account split. A planned platform called Kimi Hosted Agent will provide isolated environments for long-running tasks, and its waitlist is open.

The company also announced a developer tier with rate limits of 100 requests per minute for the API, up from 30 for K2.6. Enterprise customers can negotiate custom rate limits and dedicated compute instances. Kimi says it plans to release a smaller, distilled version of K3 later this year for edge deployment.

Related on Neura Market

More from Neura News

Product Launch

Acer Unveils Veriton RI110 Mini Workstation for Local Agentic AI

Acer unveiled the Veriton RI110 AI Mini Workstation on September 2, 2026, in Berlin. This compact desktop, featuring an Intel Core Ultra X7 processor and Intel Arc B390 graphics, supports local inference of AI models up to 120 billion parameters. It is designed for hybrid agentic AI workloads, combining local processing with cloud resources, and includes the Qubi Claw software suite for secure, autonomous AI tasks. The system offers up to 96 GB of LPDDR5X memory, 4 TB of SSD storage, and extensive connectivity options including OCuLink, Wi-Fi 7, and dual LAN ports. Availability begins in North America in Q4 2026 and EMEA in Q1 2027.

Sep 2·4 min read