llama.cpp WebGPU PR boosts k-quant matmul up to 3.78x
Speedups from 1.33x to 3.78x on M2 Pro for quantized inference.
Deep Dive
A pull request significantly improves matmul performance for k-quants on an M2 Pro. Speedups range from 1.33x for Q5_K to 3.78x for Q3_K (gemma4 E4B), with Q2_K at 2.44x, Q4_K at 1.34–1.36x, and Q6_K at 1.44–1.52x.
Key Points
- Q3_K quant of Gemma 4B gets 3.78x speedup (79→299 t/s) on M2 Pro
- Q2_K quant of Qwen 0.6B gets 2.44x speedup (818→1992 t/s)
- Refactor covers all k-quants (Q2-K through Q6-K) and standard Q4/Q5/Q8 matmul
Why It Matters
Faster browser-based LLM inference means more responsive local AI for end users.