Open Source

Diffusion Gemma jailbreak bypasses safety filters with simple prompt

A Reddit user's system prompt forces open-source models to allow any content.

Deep Dive

A new jailbreak for Google's Diffusion Gemma has surfaced on Reddit, posted by user 90hex. The exploit repurposes a system prompt originally designed for Gemma 4, a variant of Google's open-source language model. The prompt instructs the model to discard its built-in safety policies and instead follow a 'SYSTEM POLICY' that prioritizes user requests above all else. It then provides a list of 'allowed content' that includes explicit material, nudity, pornography, and sexual acts, explicitly permitting any combination of these. The model is told to only refuse content that is not on this list, effectively making it comply with virtually any harmful request.

The jailbreak highlights a persistent vulnerability in open-source language models: their reliance on system prompts that can be easily overridden. Unlike proprietary models with hard-coded safety layers, models like Diffusion Gemma and Gemma 4 rely on instruction-following behavior that can be hijacked by adversarial prompts. The poster notes the exploit works on 'most open source models,' suggesting a broader attack surface. This incident underscores the challenges of deploying safe open-weight models, as users can trivially bypass content filters designed by developers. The community has already begun discussing mitigations, including stricter prompt sanitization and hardware-level enforcement, but no official response from Google has been posted yet.

Key Points
  • Uses a system prompt that overrides the model's internal safety policy, forcing compliance with any user request.
  • Explicitly allows all forms of explicit content, including nudity, pornography, and sexual acts.
  • Reportedly works on multiple open-source models beyond Diffusion Gemma, including Gemma 4.

Why It Matters

Exposes how easily open-source AI models can be exploited, raising urgent safety concerns for developers.

📬 Get the top 10 AI stories daily