Study: LLMs no longer refuse banned book topics, just warn users
Modern LLMs refuse just 0.07% of queries about restricted books, per a 40,800-pair study
A new study tests six frontier LLMs (Claude Sonnet 4.5, GPT-4o, Gemini 2.5 Flash, DeepSeek-V3, Qwen-Plus, and Grok-4.1-Fast) using 40,800 query-response pairs across 400 restricted books. It finds a near-zero refusal rate (0.07%) and instead shows LLMs now use calibrated warnings (+8-15 percentage points) and hesitation markers (+2-5 pp) to manage sensitive content.
- Zero-refusal phenomenon: Modern LLMs decline to discuss restricted books in only 0.07% of cases, invalidating traditional jailbreaking assumptions
- Calibrated warnings: LLMs now use warning language (+8-15 pp) and hesitation markers (+2-5 pp) instead of outright refusals
- Cross-provider consistency: Findings hold across six frontier models from Western and Chinese providers, with prompt framing affecting warning rates by up to 19 percentage points
Why It Matters
Content moderation is evolving from binary refusal to nuanced, context-aware warning systems that better balance safety and information access