Anthropic's Fable guardrails anger researchers with overzealous blocks
Even asking for a code review triggers safety blocks on Anthropic's new model.
Anthropic launched Fable on Tuesday as a more accessible version of its powerful cybersecurity model Mythos, but the new model's guardrails are drawing sharp criticism from security researchers. The safety measures reject any prompt that touches on cybersecurity topics, even benign requests like reading a blog post or asking for a code review. Valentina Palmiotti of IBM X-Force noted that the model shuts down and claims its 'safety measures flagged this message for cybersecurity or biology topics.' Veteran researcher Matt Suiche described the restrictions as keyword-based, where anything in the lexical field of 'cybersecurity' triggers the guardrail and causes Fable to fall back to Claude Opus 4.8. He acknowledged it's better to over-catch initially and relax later.
Mythos, originally released in April under Project Glasswing, was restricted to select organizations for securing critical infrastructure. Last week, Anthropic expanded access to hundreds of organizations across 15 countries. For researchers who want fewer limitations, Anthropic runs a Cyber Verification Program where approved applicants get reduced guardrails. OpenAI has a similar Trusted Access for Cyber program. Despite the intent to prevent misuse for malware or bioweapons, many professionals find the current approach too blunt. Anthropic has not yet commented on the backlash, but experts expect the guardrails to evolve with more collaboration between frontier model companies and the cybersecurity industry.
- Anthropic's Fable model blocks even innocuous cybersecurity terms like 'code review' or 'blog post', forcing fallback to Claude Opus 4.8.
- Researchers criticize the keyword-based guardrails as too aggressive, but acknowledge over-capturing is better than under-capturing in early releases.
- Anthropic's Cyber Verification Program and OpenAI's Trusted Access offer reduced restrictions for vetted cybersecurity professionals.
Why It Matters
Overly restrictive AI guardrails could hinder legitimate security research, slowing innovation in real-world cyber defense.