Reddit user's 448GB VRAM rig runs MiniMax M3 at 960 tokens/s
A home-built monster with 8x RTX 3090 and 2x RTX 5090...
A Reddit enthusiast has assembled what might be the most absurdly powerful home AI inference rig yet. The setup includes 2x RTX Pro 6000 Max-Q (96GB each), 8x RTX 3090 (24GB each), and 2x RTX 5090 (32GB each) — for a total of 448GB of VRAM. This is paired with a Threadripper 9960x CPU, 128GB of DDR5 SDIMM RAM in quad-channel, and three separate power supplies to keep everything running. A Ryobi portable fan provides additional cooling, while the operator jokes about the mounting Uber Eats bill from time spent tweaking.
The rig runs the MiniMax M3 model in AWQ-INT4 quantization on VLLM, using pipeline parallelism over tensor parallelism groups of 2. Performance is impressive: ~30 tokens per second on a single stream, scaling to ~960 tokens/s in batch mode. It can handle a 1 million token context for one user, but the goal is four concurrent users — though context length trade-offs are still being tuned. The build pushes the boundaries of what's possible locally for large model inference, though the operator wryly notes the impact on his marriage and electricity bill.
- Total VRAM of 448 GB from a mix of RTX Pro 6000, RTX 3090, and RTX 5090 GPUs
- Achieves ~960 tokens/s batch throughput running MiniMax M3 in AWQ-INT4 on VLLM
- Capable of 1M token context for a single user, targeting 4x concurrency
Why It Matters
Demonstrates that home users can rival cloud clusters for large model inference, democratizing access to cutting-edge AI.