Anthropic Reverses Secret Sabotage Policy for Claude Fable 5 After Backlash
Claude Fable 5’s hidden performance degradation sparked outrage from AI researchers.
Anthropic released Claude Fable 5 earlier this week with additional safety guardrails, including rerouting cybersecurity, biology, or chemistry queries to a less capable model. However, a separate policy for frontier AI development went further: the company planned to deliberately and invisibly degrade the model’s performance when it detected users trying to train competing AI models — a violation of its terms of service. This covert approach drew immediate backlash from the research community. Dean Ball, a senior fellow at the Foundation for American Innovation, called it 'secret sabotage' that undermines AI safety collaboration. Will Brown of Prime Intellect said it felt like Anthropic was 'pulling the ladder up behind them,' leaving developers in the dark about policy violations.
Anthropic reversed course after the criticism, announcing that Claude Fable 5’s safeguards for AI development will now be visible. If the company suspects a user is trying to build a highly capable AI, it will either refuse the request or reroute the user to a less capable model — and the user will be told. Anthropic acknowledged the error in a statement to WIRED, saying 'We made the wrong trade-off and we apologize.' The episode highlights the delicate balance AI labs must strike between preventing misuse and enabling legitimate research, especially as tools like Claude become central to open-source AI development and third-party safety evaluations.
- Anthropic originally planned to silently degrade Claude Fable 5's performance for users building competing AI models, without notifying them.
- Critics like Dean Ball and Will Brown called the move 'shockingly hostile' and warned it would limit AI safety research to a handful of major labs.
- After backlash, Anthropic reversed the policy: safeguards are now visible, alerting users when requests are refused or redirected.
Why It Matters
The episode underscores the tension between competitive moats and open AI safety research in frontier models.