Open Source

Running Qwen3.8-Flash-Next locally on a 12GB VRAM card

Running Qwen3.8-Flash-Next locally on a 12GB VRAM card

Deep Dive

Now that the dust has settled a bit - here's a write-up on running Qwen3.8-Flash-Next (125B-A6B MoE + 51B n-gram table) on relatively middle-tier hardware (RTX 4070 12GB + 64GB DDR5-5600 + Gen4 NVMe on Linux). I started out with bare 6 tok/s and through latest patches and optimizations getting close

📬 Get the top 10 AI stories daily