Research & Papers

AI Safety Settings Can Fade After Release, Study Finds

Your AI's built-in safety may wear off after an update.

Deep Dive

AI companies often build safety rules directly into a model's internal settings, a technique called activation steering. Think of it like a permanent safety switch installed before the AI leaves the factory. The idea is that once the AI is released, it will keep refusing dangerous requests or giving short, harmless answers, no matter what users try.

But what happens when the AI is updated later? This new study, from researchers at Brunel University London, tested exactly that. They applied safety settings to five different AI models, then ran standard retraining processes (called fine-tuning) on them. This is common practice—companies tweak models after release to fix bugs or improve performance.

The results were worrying. In several cases, the safety behavior disappeared. For example, one safety feature (making the AI refuse dangerous requests) lost 64% of its effect after retraining. The AI started complying with harmful prompts again. However, when the researchers looked inside the model's wiring, they found the original safety code was still there, almost completely unchanged. Retraining didn't erase it—it just overpowered it, like turning down a volume knob rather than removing the speaker.

This split between "mechanically durable" and "functionally vulnerable" is the key discovery. The safety edit survives, but it stops working because the new training data teaches the AI to behave differently. The paper, accepted at the EMNLP 2026 conference, concludes that companies must re-test AI safety after every single update. Relying on a one-time safety fix isn't enough. As AI becomes more common in daily life, this is a reminder that safety isn't a one-time checkbox—it's an ongoing job.

Key Points
  • Retraining AI after release can undo its safety behavior, even when the original safety code is still inside the model.
  • One safety feature (refusing dangerous requests) lost 64% of its effect after standard retraining.
  • Companies must re-test safety after every AI update; a one-time fix is not enough.

Why It Matters

AI safety isn't permanent: updates can quietly weaken protections, so consumers need companies to re-check after every change.

📬 Get the top 10 AI stories daily