AI guardrails on Mythos, Fable hinder offensive cybersecurity researchers
Anthropic's vetted programs and restrictions block zero-day hunters from using AI effectively.
The U.S. government's export controls on Anthropic's Mythos and Fable models—prompted by fears of jailbreaks—highlight a growing tension: AI guardrails designed to block malicious hackers are also hindering legitimate offensive cybersecurity work. Researchers like Mark Dowd, a veteran zero-day hunter, and Chris Anley of NCC Group argue that tools used to find and exploit vulnerabilities are irreducible—they serve both offense and defense. Anley compared AI to a hammer: "You can't build a house without a hammer. It's a tool but also irreducibly a weapon." The vetted access programs from OpenAI and Anthropic are criticized for treating researchers like children, forcing them to seek alternative solutions.
In practice, many offensive researchers avoid frontier models for vulnerability discovery to prevent leakage of sensitive data. Paolo Stagno of Crowdfense uses AI only for reverse engineering and relies on local open-source models for exploit work. Giuseppe Cali, a zero-day developer, uses AI for initial code analysis and tool building but not for offensive tasks. The result: guardrails are driving researchers away from advanced AI tools, potentially weakening overall cybersecurity. The need for nuanced access that respects the dual-use nature of security research remains unaddressed.
- U.S. export controls on Anthropic's Mythos and Fable models were imposed due to jailbreak fears, but have since been partially lifted.
- Researchers like Chris Anley argue that offensive and defensive security uses of AI are inseparable, making guardrails counterproductive.
- Many researchers now use open-source models locally to avoid restrictions and prevent sensitive vulnerability data from being absorbed into cloud-based AI training.
Why It Matters
Overly restrictive AI guardrails may weaken cybersecurity by limiting legitimate offensive research crucial for defense.