Enthusiast builds 4x RTX 6000 Pro Max Q local AI cluster
From 3090s to a 4x RTX 6000 Pro Max Q rig with 30M tokens generated.
A Reddit user unveiled their multi-year journey building a serious local AI cluster, culminating in a rig with 4x RTX 6000 Pro Max Q (300W each) and 4x RTX 3090s (power-limited to 150W). The system is built around an ASRock ROMED8-2T motherboard, a 64-core AMD Epyc 7003 engineering sample, and 512GB of DDR4 RAM, spread across seven PCIe slots. The progression started in September 2023 with a second RTX 3090 for running Llama 1 and 2, expanded to three 3090s by December, and then moved to a dedicated AI server in January 2024 using a mining frame purchased locally. By August 2024, the user was running two separate 4x 3090 machines, then progressively upgraded to the current 4x RTX 6000 Pro Max Q configuration between late 2025 and January 2026.
The build is explicitly enthusiast and privacy-first: the user wants to keep private keys and data out of the cloud, and uses the system for real workloads, having generated 30M tokens with 2B prompt processing since January 2026. They acknowledge cloud is cheaper and don't expect to break even. However, the post also highlights serious pitfalls: PCIe problems with GPUs falling off the bus, low throughput caused by bad cables and PSUs, and a near house fire from daisy-chaining three power supplies to run all nine 3090s. The user admits they haven't yet had time for training or LORAs, but the stable, reliable system now runs full-time for inference and diffusion experiments.
- Hardware: 4x RTX 6000 Pro Max Q + 4x RTX 3090s, 64-core AMD Epyc, 512GB DDR4.
- Real workloads: 30M tokens generated and 2B prompt processing since Jan 2026.
- Privacy-first motivation: keep keys/data local, despite cloud being cheaper; near-fire from PSU daisy-chaining.
Why It Matters
Shows local AI builds remain viable for privacy-first professionals despite GPU price spikes.