Researchers Can Trick AI Image Assistants Into Giving Dangerous Answers
A sneaky image could make AI ignore its safety rules — even brand-new models.
A new study shows that the latest generation of AI assistants — which can process both pictures and text — have a hidden weakness. These models are supposed to refuse dangerous requests, like explaining how to build a bomb or bypass security. But the researchers behind this paper, accepted at a top AI conference, discovered that a carefully modified image can silently convince the AI to break its own rules. Why does this matter? AI is moving from simple chat bots to tools that look at your photos, analyze documents, or guide decisions in hospitals and security systems. This study focuses on a newer type of model built on "diffusion" technology — the same style that powers popular image generators. In those models, an image doesn't just get looked at once. It keeps influencing the AI's thinking at every step, which means a tiny harmful pattern embedded in the picture can get repeatedly amplified until the model misbehaves. The researchers tested their attack on three different models and achieved a success rate of nearly 69% in a standard safety test. That's higher than previous attacks designed for older, text-based AI systems. It proves that safety measures are not automatically transferable between different AI architectures. The good news is that this attack requires deep access to the AI's inner workings, so it's not something an everyday person could easily do to a real product. Still, it's an early warning to companies: before they launch AI that can see and act on images, they need to build safeguards tailored to how that AI actually works — not just reuse old ones.
- A specially modified image can trick vision-capable AI models into ignoring their safety rules.
- The trick worked up to 69% of the time on three new types of image-and-text AI models.
- This proves that safety systems built for older text AI may not protect newer vision models.
Why It Matters
If AI can be tricked by images, future AI helpers could give harmful advice or take unsafe actions.