Google's Gemma 4 update fixes tool calling, slashes laziness, adds Flash Attention 4
New chat templates boost reliability, cut refusal, and accelerate inference on Hopper GPUs.
Deep Dive
Key Points
- Updated chat templates improve reliability of tool calling (function calling) and reduce model refusal or early termination.
- Flash Attention 4 is now supported on Hopper GPUs (H100/H200), enabling faster inference with lower memory footprint.
- New interactive Hugging Face guide helps users tune Gemma 4’s vision token budget for optimal image-context trade-offs.
Why It Matters
More reliable tool calling and faster inference make Gemma 4 a stronger choice for production-grade AI agents.