Viral Wire

Google DeepMind unveils DiffusionGemma 26B at 1000 tokens/sec, Gemini 3.5 Live Translate

New open model generates text 4x faster than prior diffusion methods.

Deep Dive

Google DeepMind revealed a suite of initiatives spanning safety, efficiency, and new capabilities. The centerpiece is DiffusionGemma 26B, an open-weight model released June 10, 2026 that achieves 4x faster token generation through parallel text diffusion — hitting over 1,000 tokens per second on a single H100 GPU. This marks a significant leap over standard autoregressive models by generating multiple tokens simultaneously.

Alongside, the company introduced Gemini 3.5 Live Translate, enabling fluid, natural voice translation in real time. They also announced Gemma 4 12B, a new unified, encoder-free multimodal model that processes images and text without a separate vision encoder. Finally, DeepMind committed to increased investment in multi-agent AI safety research, addressing risks from systems of interacting agents.

Key Points
  • DiffusionGemma 26B is open (weights available) and runs 4x faster than prior diffusion models, achieving 1000+ tokens/sec on one H100.
  • Gemini 3.5 Live Translate provides natural-sounding voice translation without noticeable delay.
  • Gemma 4 12B is an encoder-free multimodal model, simplifying architecture for vision-language tasks.

Why It Matters

Faster open models and real-time translation lower costs and expand access to advanced AI capabilities.

📬 Get the top 10 AI stories daily