{ "title": "Mira Murati's Thinking Machines Lab Launches Inkling, a 975-Billion-Parameter Open-Source AI Model", "body": "Mira Murati’s AI startup, Thinking Machines Lab, released its first production-ready model on Thursday. Inkling, a 975-billion-parameter open-weights multimodal model, tops the charts for US open-source AI but falls behind Chinese competitors in overall performance and factual accuracy.\n\nThe company acknowledged the gap in its announcement. “Inkling is not the strongest overall model available today,” Thinking Machines stated. The model is designed as a flexible base for customization, not a one-size-fits-all champion.\n\n## A Mixture-of-Experts Giant\n\nInkling uses a Mixture-of-Experts Transformer architecture. It has 975 billion total parameters, with 41 billion active at any time. This design lets the model stay efficient while handling a wide range of tasks.\n\nThe model natively processes text, images, and audio. It supports a context window of up to one million tokens. Thinking Machines pre-trained Inkling on 45 trillion tokens of public and synthetic data, including text, images, audio recordings, and videos. The company noted that the training set includes public data that “may be subject to intellectual property protection.”\n\nTo generate synthetic training data, Thinking Machines used Kimi K2.5, a Chinese model that also serves as the basis for the coding tool Cursor’s model.\n\n## Benchmark Performance: Leading the US, Chasing China\n\nOn the Artificial Analysis Intelligence Index, Inkling scored 41. That makes it the top US open-weights model on that benchmark. The previous US leader, Nemotron 3 Ultra, scored 38. Other US models trailed further: Gemma 4 31B scored 29, and gpt-oss-120b scored 24.\n\nBut Chinese open-source models still outperform Inkling overall. On the GDPval-AA v2 agent-based benchmark, Inkling reached an Elo rating of 1,238. Chinese models Kimi K2.6 scored 1,190, and DeepSeek v4 Flash max scored 1,189. Inkling’s lead is narrow.\n\nOn the Tau-3 banking benchmark, Inkling scored 24%. DeepSeek v4 Flash max scored 23%, and Kimi K2.6 scored 21%. Inkling edges ahead in this domain-specific test.\n\n## Factual Accuracy: A Weak Spot\n\nInkling’s factual accuracy results are likely to limit its use in applications needing highly reliable information. On the AA Omniscience benchmark, Inkling scored +2. Nemotron 3 Ultra scored -1. But Inkling’s accuracy was only 40%, and its hallucination rate was 63%.\n\nThat high hallucination rate means the model frequently generates incorrect information. For tasks requiring precise answers, users may need to apply additional safeguards or fine-tuning.\n\n## Pricing, Efficiency, and a Surprising Compact Model\n\nInkling costs $1.87 per million input tokens and $4.68 per million output tokens with a 64K context window. For context windows up to 256,000 tokens, the price rises: input costs $3.74 per million tokens, cached input costs $0.748 per million tokens, and output costs $9.36 per million tokens.\n\nInkling is efficient with tokens. It averages 25,000 output tokens per Intelligence Index task. By comparison, GLM-5.2 max uses 43,000 tokens, Kimi K2.6 uses about 38,000 tokens, and DeepSeek v4 Pro max uses about 37,000 tokens on the same tasks. This efficiency could lower costs for users running many queries.\n\nThe model offers continuously adjustable “thinking effort,” letting users trade speed for deeper reasoning.\n\nThinking Machines also previewed Inkling-Small, a compact version with 276 billion total parameters and 12 billion active. On the GPQA Diamond benchmark, Inkling-Small scored 88.3%, beating the full Inkling’s 87.2%. On the HLE benchmark with tools, Inkling-Small scored 46.6%, again slightly ahead of Inkling’s 46.0%.\n\nThe company credits changes to pre-training data and the training process for these results. Thinking Machines plans to publish full Inkling-Small weights once testing is complete.\n\n## Access and Competitive Landscape\n\nInkling’s weights are freely available on Hugging Face. Users can also access the model through Tinker, Thinking Machines’ platform for adapting AI models to specific tasks. The company expects multimodal support, efficient processing, and fine-tuning options to set the model apart from competitors.\n\nInkling enters a crowded field. US open-weights models now have a new leader, but Chinese models like Kimi K2.6 and DeepSeek v4 still hold the edge in overall performance. Inkling costs slightly more than open-source Chinese models like GLM-5.2 and DeepSeek v4, but its lower token usage may offset that difference for some workloads.\n\nMira Murati, former OpenAI CTO who played a key role in developing ChatGPT, founded Thinking Machines Lab. Inkling is the company’s first major release. The model’s high hallucination rate and low factual accuracy suggest it will need careful handling in production, but its strong performance on agent-based and banking benchmarks shows promise for specialized applications.\n\n## Related on Neura Market\n\n- Thinking Machines Lab Company Profile\n- Open-Weights AI Models Comparison\n- Artificial Analysis Intelligence Index" }
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.

