Open Source

llama.cpp WebGPU PR boosts k-quant matmul up to 3.78x

Speedups from 1.33x to 3.78x on M2 Pro for quantized inference.

Deep Dive

A pull request significantly improves matmul performance for k-quants on an M2 Pro. Speedups range from 1.33x for Q5_K to 3.78x for Q3_K (gemma4 E4B), with Q2_K at 2.44x, Q4_K at 1.34–1.36x, and Q6_K at 1.44–1.52x.

Key Points
  • Q3_K quant of Gemma 4B gets 3.78x speedup (79→299 t/s) on M2 Pro
  • Q2_K quant of Qwen 0.6B gets 2.44x speedup (818→1992 t/s)
  • Refactor covers all k-quants (Q2-K through Q6-K) and standard Q4/Q5/Q8 matmul

Why It Matters

Faster browser-based LLM inference means more responsive local AI for end users.

📬 Get the top 10 AI stories daily