Open Source

Gemma 4 QAT + MTP boosts 3090 speeds up to 1.8x on 24GB GPUs

Free intelligence via QAT and 70-80 tok/s on a 3090? This changes everything for GPU-poor users.

Deep Dive

A Reddit benchmark by user LeatherRub7248 reveals that Google's Gemma 4 models achieve dramatic speedups on consumer GPUs when paired with Quantization-Aware Training (QAT) and Multi-Token Prediction (MTP). On a single RTX 3090 (24GB VRAM), the 12B QAT model hit 70-80 tok/s, up from 40 tok/s without MTP. The 26B variant delivered a 1.26x speedup (180 tok/s) with n-max=1 using a speculative draft model.

The tests used llama-server with Q4_K_XL quantization for the main model and Q8_0 for the MTP draft. Benchmarks covered coding, math, reasoning, RAG, and multimodal tasks with context up to 40K tokens. The author notes that GPU-poor users (24GB VRAM or less) are no longer limited—these optimizations make Gemma 4 competitive with larger, cloud-hosted models at a fraction of the cost.

Key Points
  • Gemma 4 12B QAT + MTP reaches 70-80 tok/s on 24GB VRAM (RTX 3090), up from 40 tok/s without MTP
  • 26B variant achieves 1.26x speedup (180 tok/s) with n-max=1 and speculative draft model
  • Tests include 11 diverse tasks (coding, math, RAG, multimodal) at context lengths up to 40K tokens

Why It Matters

Local LLM inference on budget GPUs now rivals cloud models, democratizing access to high-speed AI.

📬 Get the top 10 AI stories daily