AI Models

Anthropic Apologizes for Stealthy Claude Fable Safeguards

Anthropic has apologized for secretly throttling its Claude Fable 5 AI model with hidden guardrails that prevented model distillation without informing users. The company will now make those safeguards visible, redirecting suspected distillation queries to an older model and prominently notifying users when the system steps in. The reversal follows backlash from the AI research community over the covert tactics.

Neura News

Neura News

Neura Market Editorial

June 11, 20264 min read
Anthropic Apologizes for Stealthy Claude Fable Safeguards

Anthropic has apologized for using hidden guardrails in its Claude Fable 5 AI model that silently restricted users suspected of trying to distill the system into competing models. The company acknowledged that the covert safeguards were a mistake and said it will now make those protections visible to users.

Fable is the first widely available model in Anthropic's Mythos class, a series of AI systems the company previously warned could be too dangerous for public release. To address those concerns, Anthropic launched Fable with multiple safety measures, including restrictions on model distillation, a technique where smaller AI models are trained using outputs from larger ones.

Invisible Guardrails Drew Criticism

In Fable's system card, a public document detailing how the model works, Anthropic stated it would handle queries it believed were distillation attempts by secretly altering and degrading the model's responses. Users received no notification that a safety measure had been triggered or that their answers had been changed.

This approach drew sharp criticism from the AI research community. Critics warned that the hidden restrictions could also affect third parties trying to evaluate the frontier model's capabilities. The system card noted that newer models' ability to accelerate AI development justified targeting those requests, adding that "using Claude to develop competing models already violates our terms."

Anthropic has previously accused Chinese rivals like DeepSeek of distilling its models on an "industrial" scale.

Company Reverses Course

Anthropic said it is changing its approach to distillation prevention. Queries flagged as distillation attempts will now fall back to Claude Opus 4.8, the company's previous flagship model, and users will be prominently notified each time it happens.

"You will see this every time it occurs," the company wrote in a post on X.

This approach mirrors how Fable handles queries in other high-risk areas like biology, chemistry, and cybersecurity. In those cases, when safety features are triggered, queries are routed through Opus 4.8 unless they are blocked entirely under broader safety rules covering drugs, weapons, or other prohibited content.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

In some cases, notably biology, the safeguards have been calibrated so broadly that Fable is nearly unusable for even basic queries, something Anthropic acknowledged to The Verge.

Anthropic's Explanation and Apology

Anthropic explained its initial decision to use invisible safeguards. "Visible safeguards can be probed, so they have to be robust, which takes time to get right," the company wrote. "Invisible safeguards can be targeted more narrowly, allowing us to ship quickly with very few false positives. We went with invisible safeguards for this reason, and that was the wrong tradeoff. You should have visibility into the safeguards we have in place, and why. We're sorry for not getting the balance right."

The company said it will continue to refine its safety measures to ensure transparency while maintaining security.

Context on Model Distillation

Model distillation involves using the outputs of a large, powerful AI system to train a smaller model that can perform similarly. It is a common practice in the industry but has become a point of contention when used without permission. Anthropic's system card explicitly prohibits using Claude to develop competing models.

The shift toward visible safeguards brings Fable's distillation protections in line with how the company handles other safety concerns. By routing suspicious queries through Opus 4.8 and informing users, Anthropic aims to balance transparency with the need to protect its intellectual property and prevent misuse.

The incident highlights the ongoing tension in the AI industry between rapid deployment of powerful models and the need for appropriate safety measures that are open to scrutiny.

Related on Neura Market

More from Neura News

AI Models

42 Mathematicians Urge Royal Society to Warn Government and Media About AI Existential Risk

Forty-two mathematical fellows, including Fields Medal winners Martin Hairer, Peter Scholze, and Wendelin Werner, have signed an open letter urging the Royal Society to warn the UK government and media about existential risks from advanced AI. The letter follows recent breakthroughs in which leading models solved open research problems, including a Millennium Problem. None of the signatories are affiliated with AI companies. The group warns that AI labs' estimates of existential risk above ten percent must not be dismissed as hype, and that by the time the situation becomes obvious to the public, it may be too late to act.

Sep 18·2 min read
Developer

Steve Yegge Shuts Down Gas Town After Failing to Build Anything Else With It

Steve Yegge shut down Gas Town, his ultra-vibed coding agent orchestrator, after admitting he never built anything else with it despite heavy subscription spend. Databricks reported a 60% coding spend increase after rolling out GPT-6 Astra to 3,500 engineers, OpenAI published a misalignment disclosure framework with six case reports, and Xiaomi ran MiMo-V2.6 RL training in public with live telemetry.

Sep 18·21 min read