Developer Tools

Anthropic's Fable 5 refuses to discuss cybersecurity and biology topics

New Claude model blocks dangerous queries but may frustrate harmless users.

Deep Dive

Anthropic launched Claude Fable 5, a 'Mythos-class' model that the company says surpasses its previous Opus models. It uses classifiers to refuse or redirect queries on cybersecurity, biology, and chemistry to the earlier Claude Opus 4.8 model. Over 1,000 hours of red-team testing found no universal jailbreaks for Fable 5. According to the article, Mythos 5 scored 78 percent on the ExploitBench test, up from 40 percent for Opus 4.8 and 69 percent for Mythos Preview. Anthropic tuned safeguards to be "stricter than ideal," causing false positives in less than five percent of testing sessions.

Key Points
  • Fable 5 scores 78% on ExploitBench (cybersecurity exploits), up from Opus 4.8's 40%.
  • Over 1,000 hours of red-team testing found zero universal jailbreaks.
  • False positives (harmless queries blocked) occur in less than 5% of sessions.

Why It Matters

Balancing powerful AI capabilities with safety — Anthropic deliberately limits access to prevent misuse by malicious actors.

📬 Get the top 10 AI stories daily