Unsloth releases MTP GGUF weights for Gemma 4 (12B–31B)
Multi-token prediction weights cut inference costs in half for local AI.
Deep Dive
Unsloth has pushed MTP GGUF weights in Q8, F16, and BF16 formats for Gemma 4 variants: 31B, 26B-A4B, and 12B.
Key Points
- Three Gemma 4 variants supported: 12B, 26B-A4B, 31B – all with MTP decoding for faster inference.
- Quantizations offered: Q8 (8-bit), F16 (half-precision), BF16 (bfloat16) – choose based on hardware/quality trade-off.
- Up to 2x lower VRAM usage: 31B model fits on 16 GB GPUs vs previous 24 GB barrier.
Why It Matters
Unsloth makes Gemma 4’s multi-token prediction practical on consumer hardware, slashing costs and latency for local AI.