Open Source

Google's Diffusion Gemma is 4x faster but makes 6x more factual errors

Fast text generation comes at a steep accuracy cost in new diffusion model.

Deep Dive

Google's Diffusion Gemma is faster than standard Gemma 4 (763 vs 218 tok/s) but far less accurate. In a benchmark on a single H100 (FP8) across three tasks—a Steve Jobs biography, a Tetris history, and a BeOS story—Gemma 4 got 45 facts right and 5 wrong, while Diffusion Gemma got 33 right and 28 wrong. Mistakes included inventing a mother named Clara Clley for Jobs, a nonexistent colleague for Tetris's creator, and pricing the BeBox at $9,999 (real cost: $1,600). Errors increased as topics got less popular: 4 on Jobs, 12 on Tetris, 12 on BeOS. The reason: Diffusion Gemma outputs 256 tokens at once and refines them for smoothness, not accuracy. As Google noted, "quality is lower, use regular Gemma 4 when facts matter."

Key Points
  • DiffusionGemma achieves 763 tok/s vs 218 tok/s for Gemma 4 (4x faster) on a single H100 in FP8.
  • Fact accuracy: DiffusionGemma scored 33/61 correct (54%) while Gemma 4 scored 45/50 correct (90%).
  • Errors increased with topic obscurity: 4 mistakes on Steve Jobs, 12 each on Tetris and BeOS.
  • Diffusion model fabricated names (Clara Clley, Geri Gulovik) and wrong prices ($9,999 vs $1,600).

Why It Matters

Trade-off between speed and accuracy: diffusion models excel for creative writing but fail when factual precision is critical.

📬 Get the top 10 AI stories daily