Qwen 3.5 122B model tops coding tests on 64GB VRAM
30 tok/s with 100k context on a single GPU—no cloud needed.
Deep Dive
Jorlen reports using an unsloth Qwen 3.5 122b-a10b (UD-IQ4_NL) with a 100k bf16 context window. Some layers offload to CPU/RAM, achieving around 30 tok/sec. After hours of testing, they are deeply impressed and think this model may become their daily driver. They also use Qwen 3.6 models and ask what others with similar VRAM capacity are using.
Key Points
- Runs on 64GB VRAM using Unsloth IQ4_NL quantization.
- Delivers 30 tok/s with a 100k bf16 context window.
- Outperforms alternatives tested, becoming a daily coding driver.
Why It Matters
Developers with 64GB GPUs can now run top-tier coding models locally with minimal latency and full privacy.