Developer Tools

Ollama v0.31.1 boosts Gemma 4 speed 90% on Apple Silicon

Gemma 4 runs nearly 2x faster on Mac with zero configuration.

Deep Dive

Ollama's v0.31.1 delivers a major performance boost for Gemma 4 on Apple Silicon. The update uses multi-token prediction (MTP) to draft multiple tokens simultaneously, achieving nearly 90% faster token generation on average. This speedup is fully automatic—Ollama tunes how many tokens to draft on the fly, so no manual configuration is needed. The improvement is most notable on coding-agent benchmarks, making local AI development far more responsive.

Beyond the headline speed increase, v0.31.1 tightens Gemma 4 MoE model loading in the MLX engine, includes a new small-batch matmul kernel, and updates the underlying llama.cpp engine to build 9840. These changes improve overall stability and performance for all models, not just Gemma 4. The combination of auto-tuned MTP and engine upgrades means users get faster inference with zero friction—ideal for developers running AI models locally on Mac hardware.

Key Points
  • Gemma 4 token generation is nearly 90% faster on Apple Silicon via multi-token prediction (MTP)
  • Speedup is auto-tuned at runtime with no configuration required and no output changes
  • Includes MLX engine updates and llama.cpp build 9840 for broader performance improvements

Why It Matters

Ollama makes Gemma 4 practical on Mac for developers, enabling faster local AI workflows.

📬 Get the top 10 AI stories daily