Open Source

Qwen 3.5 122B model tops coding tests on 64GB VRAM

30 tok/s with 100k context on a single GPU—no cloud needed.

Deep Dive

Jorlen reports using an unsloth Qwen 3.5 122b-a10b (UD-IQ4_NL) with a 100k bf16 context window. Some layers offload to CPU/RAM, achieving around 30 tok/sec. After hours of testing, they are deeply impressed and think this model may become their daily driver. They also use Qwen 3.6 models and ask what others with similar VRAM capacity are using.

Key Points
  • Runs on 64GB VRAM using Unsloth IQ4_NL quantization.
  • Delivers 30 tok/s with a 100k bf16 context window.
  • Outperforms alternatives tested, becoming a daily coding driver.

Why It Matters

Developers with 64GB GPUs can now run top-tier coding models locally with minimal latency and full privacy.

📬 Get the top 10 AI stories daily