Reddit demands 80–160B AI models for high-RAM, slow-bandwidth devices
Users with 96GB+ unified memory can't run modern frontier models efficiently.
A viral Reddit post from user Storge2 highlights a growing gap in the open-source AI model landscape: hardware with large unified memory (96 GB to 128 GB) but relatively slow memory bandwidth. Affected devices include Apple Macs with M-series chips (≥96 GB), AMD Ryzen AI 395 laptops, NVIDIA DGX Spark, and multi-GPU setups like four RTX 3090s (192 GB VRAM) or even DDR5 systems with 128 GB system RAM. These machines can load models in the 80–160 billion parameter range but are too memory-constrained for massive 500B+ models (e.g., Deepseek V4 Pro, Kimi 2.7, GLM 5.2) and too bandwidth-limited to run smaller, dense models efficiently at high speeds.
Recent releases have polarized: tiny models (Qwen 3.6 35B, Gemma 4 26B) waste the memory capacity, while giant models exceed VRAM/unified memory limits. The few mid-size options (Qwen 3.5 122B, Nemotron 3 120B) are already outdated. The community specifically requests models in the 80–160B range with 10B sparse activation—a design that keeps memory footprint manageable while exploiting the large available RAM. Examples desired: Qwen 3.6/3.7 122B, Gemma 4 122B, GLM 5.2 Air, or a new Deepseek V4 Mini at 100B.
This gap affects professionals running local inference for research, code generation, and enterprise workflows on high-end consumer hardware. Without such models, these powerful but bandwidth-limited devices are forced to rely on last-generation checkpoints, undermining the value of their expensive memory configurations. The post calls on model developers (Qwen, Deepseek, Google, GLM) to prioritize a mid-range sparse architecture that balances intelligence, memory footprint, and bandwidth constraints.
- Users of Apple Silicon (>96GB), AMD Ryzen AI 395, and DGX Spark cannot run modern 500B+ models due to bandwidth limits, but 27B–31B models underutilize their RAM.
- The community calls for 80–160B sparse models (e.g., 100B active, 10B sparse) to fit within 64–128 GB unified memory while leveraging available capacity.
- Existing mid-size options like Qwen 3.5 122B are outdated; newer models (GLM 5.2, Deepseek V4 Pro) are too large, leaving a critical gap in the local AI ecosystem.
Why It Matters
A missing model tier is leaving high-RAM, low-bandwidth hardware underpowered—stalling local AI for professionals.