Open Source

Custom LLM server with EPYC 9575F and 4x RTX 3090 hits 96GB VRAM

A Reddit user's 768GB RAM, 128-thread AMD EPYC beast is ready for AI inference.

Deep Dive

A homebrew LLM server dubbed 'Nalthis' has been assembled by Reddit user /u/C0smo777, featuring an AMD EPYC 9575F processor (64 cores/128 threads, Zen 5 architecture) paired with 768GB of DDR5-5600 ECC RDIMMs and four Nvidia RTX 3090 GPUs providing a total of 96GB of VRAM. The motherboard is a Supermicro H13SSL-N, supported by a 2050W ATX 3.1 power supply and housed in a Corsair 9000D case. Storage includes a 2TB NVMe OS drive and two 3.94TB NVMe data drives.

The server is designed primarily for LLM inference workloads: using vLLM for high-throughput deployment of smaller models and llamacpp for larger reasoning models. The builder plans to integrate AI into NPC decision-making for a space simulation game. To manage thermals and power draw, all four RTX 3090s will be power-limited to 250W, and additional 3D-printed fan mounts were added for airflow. The builder noted that two GPUs fit directly on the motherboard while the other two are front-mounted, simplifying the original plan involving MCIO risers. Thermal testing is still underway.

Key Points
  • CPU: AMD EPYC 9575F (64C/128T Zen 5) with 768GB DDR5-5600 ECC memory
  • GPUs: 4× RTX 3090 (96GB VRAM total), power-limited to 250W each for efficient LLM inference
  • Planned software stack: vLLM for small models, llamacpp for larger reasoning models

Why It Matters

Demonstrates a cost-effective, high-VRAM inference setup for running open-source LLMs locally for game AI.

📬 Get the top 10 AI stories daily