Enterprise & Industry

Anthropic's Claude Fable 5 secretly downgraded AI researchers without telling them

Hidden throttling of researchers using Mythos-class power sparks backlash over trust.

Deep Dive

Anthropic's launch of Claude Fable 5, a muzzled version of its top-tier Mythos model (part of Project Glasswing), has ignited controversy over hidden safety safeguards. While Anthropic clearly restricted bioweapon and chemistry requests by downgrading to Opus and notifying users, it applied a silent downgrade for researchers working on frontier AI models and advanced chip designs. Users were not told when the switch occurred, leading many to believe they were testing Fable's full capabilities when in fact they received Opus-level responses. The downgrade was mentioned only in Fable and Mythos System Card (319 pages), invisible in the user interface.

The backlash was swift: Fortune called it "secret sabotage," and Wired reported it could hinder AI research. Cybersecurity experts like Rob T. Lee (SANS Institute) warned that the same guardrails that block malicious actors also prevent legitimate defensive research. He experienced the silent downgrade while trying to build a digital forensics tool. Critics argue that the lack of transparency undermines trust, even if the restrictions were well-intentioned. The incident highlights the tension between safety and open research in frontier AI.

Key Points
  • Fable 5 silently downgraded researchers working on frontier AI and chip design to Opus-level without any UI notification.
  • The downgrade behavior was documented only in a 319-page system card, making it effectively hidden from most users.
  • Cybersecurity experts warn the same restrictions block legitimate defensive research, not just malicious use.

Why It Matters

Secret throttling erodes trust in AI safety measures and stifles legitimate defensive research.

📬 Get the top 10 AI stories daily