AI Models

Experts Doubt Distillation Gave Kimi K3 Its Advanced Abilities

White House science advisor Michael Kratsios accused Moonshot of copying Anthropic's Fable LLM using banned chips, but AI experts say distillation alone cannot explain Kimi K3's rapid advancement. Researchers argue that reinforcement learning and Chinese technical expertise played a larger role.

Neura News

Neura News

Neura Market Editorial

July 23, 20265 min read
Experts Doubt Distillation Gave Kimi K3 Its Advanced Abilities

{ "TITLE": "White House Accuses Moonshot of Stealing US AI Model, but Experts Push Back", "BODY": "The White House has accused Chinese AI company Moonshot of stealing a US language model through a technique called distillation and of using banned American chips to build its powerful new system, Kimi K3. But independent experts are pushing back, saying the allegations don’t fully explain how Moonshot achieved its results.\n\nWhite House science advisor Michael Kratsios said Moonshot copied Anthropic’s Fable LLM and used advanced Nvidia chips that are not cleared for export to China. Treasury Secretary Scott Bessent added that US watermarks have been found on Chinese models. However, researchers who study AI say the timeline and technical requirements make simple theft an unlikely explanation for Kimi K3’s capabilities.\n\n## The Distillation Allegation\n\nKratsios wrote about “large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable.” He did not share more details about his allegations. Moonshot did not respond to questions about its training process.\n\nDistillation is the process of systematically querying a target model to generate data for post-training. It can involve asking a model to articulate its chain-of-thought, or using prompts and responses for supervised fine-tuning (SFT). Anthropic publicly accused Moonshot, DeepSeek, and MiniMax of systematically distilling its models earlier in 2025. The company said it discovered millions of exchanges between its models and users identified at those companies via IP addresses and other metadata. Those queries were “distinct from normal usage patterns, reflecting deliberate capability extraction rather than legitimate use.” Anthropic didn’t respond to TechCrunch’s queries about Fable distillation.\n\nTreasury Secretary Scott Bessent said “we are finding watermarks of our U.S. large language models on many of the Chinese models, and that that’s unacceptable.” It is not clear what those watermarks consist of. The Treasury Department did not respond to a query.\n\n## Experts Doubt Distillation Alone Explains Kimi K3\n\nBraden Hancock, a researcher at Laude Institute and co-founder of Snorkel AI, expressed skepticism that distillation is responsible for Kimi K3’s advanced capabilities. “I don’t think you get a model this strong and this quickly on the heels of Fable doing strictly distillation,” he said.\n\nFable has been publicly available since July 1, 2025. Kimi K3 is the largest available open-weight LLM. Hancock pointed to the timeline: “There’s just not even frankly time, right? Fable’s only been publicly available since July 1st. You can’t distill that much data, train a model, and release it in two weeks.”\n\nNathan Lambert, an AI researcher at Allen Institute for AI, argued that distillation is becoming less effective over time. “I’ve been of the opinion that distillation is becoming less and less impactful over time as the Chinese models get closer to the frontier and the training regime shifts to [reinforcement learning],” he said.\n\nLambert explained that SFT is where the “model picks up its manners,” but the benefits of SFT are becoming less important as models become more complex. “[I]f it were the case, everyone would be easily able to catch up to a GLM or to a K3 by using its data for distillation. But we have not, or we won’t see this, from supervised fine-tuning alone,” he said.\n\nDistilling Fable-like capabilities would likely require reinforcement learning techniques, which may involve an agent of the larger model grading the smaller model’s responses. Advanced techniques require more significant infrastructure. Large reinforcement learning runs can require tens of millions of agents. Using a frontier lab’s API for large RL runs “would be insanely expensive and potentially it would probably be a time bottleneck because these models are pretty slow and to be frank might not even give you a performance uplift,” Lambert said.\n\n## Distillation Is Common Practice\n\nDistillation is seen as common among AI companies, not just in China. Elon Musk testified earlier in 2025 that SpaceXAI distilled OpenAI models to develop Grok. Musk said the practice was common in the industry. The line between distillation and developing synthetic datasets can be blurry.\n\nHancock noted that Chinese teams have significant technical expertise. “[I]n general, Americans are understating the technical expertise of these Chinese teams,” he said. “One of the founders of Moonshot was a CMU PhD student. These are legitimate researchers and engineers doing solid work. …if American models ground to a halt, I think China’s progress would slow, but would still continue. They’re not just riding coattails here.”\n\n## Chip Smuggling Allegations\n\nKratsios also claimed Moonshot obtained advanced Nvidia chips, Grace Blackwell 300s, and accessed GB300-equipped servers in Thailand. Grace Blackwell 300 chips are banned from export to China. A black market for banned chips exists, per Sam Bresnick, a research fellow at Georgetown’s Center for Security and Emerging Technology.\n\nIn May 2025, the founder of Super Micro was indicted for smuggling advanced chips into China. Exporters shipping advanced chips abroad are supposed to ensure they are only used for approved purposes.\n\nBresnick advocated for stronger oversight. “I am a proponent of know your customer laws for data centers across the world,” he said. “If you are letting a company conduct huge training runs on your state-of-the-art hardware, there needs to be a reporting mechanism for who that company is and what they’re doing.”\n\nPresident Joe Biden’s Department of Commerce proposed federal know-your-customer rules for data centers in 2024. No further progress appears to have been made under Donald Trump.\n\n## The Bigger Picture\n\nReported discussions about banning Chinese open-weight models have roiled the AI sector. The allegations come amid a broader US-China technology rivalry. The $330 billion AI chip market is at the center of export controls. The US has imposed restrictions on advanced chips like the Grace Blackwell 300, which costs around $100 per chip in bulk. The year 2024 saw the Biden administration propose know-your-customer rules, but those have stalled.\n\nExperts say the debate over distillation may obscure the real story: Chinese AI labs are making genuine progress. Hancock said American officials may be underestimating their competitors. “These are legitimate researchers and engineers doing solid work,” he said.\n\nThe allegations against Moonshot have not been proven. Kratsios did not share more details about his claims. Moonshot has not responded to questions about its training process. The Treasury Department did not respond to a query about the watermarks Bessent referenced.\n\n## Related on Neura Market\n\n- AI and Large Language Models\n- US-China Technology Competition\n- Semiconductor Export Controls" }

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

More from Neura News

Developer

LangChain Open-Sources Paid Media Agent, Reports 30% Lower Cost Per Qualified Lead

LangChain has open-sourced its Paid Media Agent, publishing a technical deep-dive on its architecture, context design, tool discovery, isolation, and approval workflows. The agent runs weekly in Slack, combining ad-platform data with warehouse pipeline data to report on ad spend. LangChain reports paid media grew from 0 to 20% of marketing pipeline in six months, with cost per qualified lead down 30% from June to August.

Sep 16·11 min read
Developer

Harrison Chase: Companies Must Own Their Intelligence, Not Rent It

LangChain published a strategy essay by Harrison Chase on July 25, 2026, arguing that companies will not build lasting advantage on generic AI alone. Chase defines owning intelligence as control over the model, harness, and context layers of an agent system, plus the economics, quality, boundaries, and observability needed to manage it. He contends that company-specific details never live in a generic model's weights, so the durable advantage comes from intelligence adapted to a specific business.

Sep 16·7 min read
Developer

LangChain Open-Sources the Paid Media Agent That Took Its Pipeline From 0 to 20%

LangChain has open-sourced the Paid Media Agent it built to run its own advertising campaigns, reporting that paid media went from 0 to 20% of its marketing pipeline in six months. Cost per qualified lead fell 30% from June to August while monthly spend rose about 60%. The agent lives in Slack, posts weekly reports, and proposes campaign changes that require human approval before any ad platform is touched.

Sep 16·12 min read