Open Source

Google's Gemma 4 update fixes tool calling, slashes laziness, adds Flash Attention 4

New chat templates boost reliability, cut refusal, and accelerate inference on Hopper GPUs.

Deep Dive

Key Points
  • Updated chat templates improve reliability of tool calling (function calling) and reduce model refusal or early termination.
  • Flash Attention 4 is now supported on Hopper GPUs (H100/H200), enabling faster inference with lower memory footprint.
  • New interactive Hugging Face guide helps users tune Gemma 4’s vision token budget for optimal image-context trade-offs.

Why It Matters

More reliable tool calling and faster inference make Gemma 4 a stronger choice for production-grade AI agents.

📬 Get the top 10 AI stories daily