Research & Papers

DumpsterCluster serves LLaMA-70B on 128 retired V100 GPUs for $22K

128 used V100 GPUs for $22K run LLaMA-70B at competitive throughput.

Deep Dive

A team led by Zeyu Cao and Ilia Shumailov at Oxford physically assembled a DumpsterCluster of 128 retired V100 GPUs sourced from secondary markets, spending just $22K total—versus $600K for a modern 8-GPU NVIDIA B200 system. By applying pipeline-parallel optimizations, they achieved LLaMA-70B inference throughput competitive with production-grade setups, demonstrating that older accelerators can serve modern large language models when scaled horizontally. The cluster ran continuously for one year to validate real-world reliability.

But the economic win comes with a hidden environmental cost. Because V100s are far less energy-efficient than Blackwell-era GPUs, the DumpsterCluster consumes significantly more electricity per token. The researchers found that second-hand systems generate roughly 4x higher carbon emissions per token for 8B-parameter models and over 40x higher for 70B models compared to current-generation hardware, under average grid conditions. That means repurposing retired GPUs is only sustainable when paired with low-carbon, inexpensive electricity. In regions with favorable energy economics, though, dumpster-diving clusters offer a practical path to expand AI capacity at a fraction of the cost—advancing affordability and energy security simultaneously.

Key Points
  • 128-GPU cluster built entirely from used V100s costs $22K vs. $600K for an 8-GPU B200 system
  • Pipeline-parallel optimizations enable LLaMA-70B inference competitive with production systems
  • Carbon emissions per token are 4x higher (8B) and 40x higher (70B) unless powered by clean, cheap electricity

Why It Matters

Retired GPUs can slash AI inference costs 27x, but only green energy makes the tradeoff sustainable.

📬 Get the top 10 AI stories daily