Anthropic's Fable 5 pulled offline over 'fix this code' 'jailbreak' – experts call it absurd
The model was taken down for fixing code; no actual jailbreak found, experts say.
In a controversial move, Anthropic's Fable 5 model was pulled offline following a White House emergency request after a reported 'jailbreak.' The incident, detailed by LessWrong's Zvi, has now been debunked by the only outside expert with access to the report, Katie Moussouris. She states the alleged exploit was simply asking Fable to 'fix this code'—a routine bug-fixing task. The researchers used open-source code with known CVEs plus deliberately planted vulnerabilities, then manually turned outputs into test scripts. No actual security breach occurred, and Fable provided no uplift over Opus 4.8 or GPT-5.5.
The decision has sparked outrage among CISOs and cyber defenders. Simon Willison noted that fixing bugs is precisely what coding models should do, and this sets a dangerous precedent. The shutdown of both Fable and an earlier model Mythos has left American companies without cutting-edge AI for cybersecurity, while prediction markets now show only a 55% chance of restoration by July 1. Zvi argues this permanently damages trust in U.S. AI leadership and the tech stack, as adversaries race ahead. The fiasco appears rooted in non-technical decision-makers conflating defensive coding with offensive cyber capabilities, threatening to ban models that help secure software.
- Fable 5 was taken offline after being asked to 'fix this code'—not a real jailbreak, per outside expert Katie Moussouris.
- The test used deliberately vulnerable open-source code and manual steps; Fable offered no uplift over other models.
- Prediction markets give only 55% chance of restoration by July 1, while U.S. cyber defenses suffer without access.
Why It Matters
Misguided AI shutdowns based on non-exploits risk hampering U.S. cybersecurity and eroding global trust in American AI.