Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second
Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second
Deep Dive
A while ago I posted 15 tok/s output and 100-120 tok/s prompt processing with the IQ3_XXS quant on a 12GB RTX 5070 using llama.cpp. Since then I built my own inference engine for this one model and this kind of PC. The same IQ3_XXS now runs at ~65 tok/s output and ~430 tok/s prompt processing , and