Open Source

Poolside's Laguna S 2.1 runs at 400-600 tok/s on 3x AMD V620 GPUs

96GB VRAM setup delivers impressive inference speeds for under $1,050 total.

Deep Dive

A Reddit user testing the Laguna model reports it fits Q4_K_M with 256K context @ F16 on their setup. They're running an HTML flight simulator test and seeing 400–600 tok/s prefill and 16–20 tok/s generation (no dflash enabled), calling it "not bad for the cost" — cards at $350 each. Their hardware: Dell PowerEdge R740 with dual Xeon Gold 6248R and 768 GB RAM.

Key Points
  • Poolside's Laguna S 2.1 runs on 3x AMD Radeon Pro V620 GPUs (96GB VRAM) for under $1,050 total GPU cost.
  • Achieves 400–600 tok/s prefill and 16–20 tok/s generation with 256K context at Q4_K_M.
  • Tested on a Dell PowerEdge R740 with dual Xeon Gold 6248R and 768GB RAM; dflash not yet enabled.

Why It Matters

Shows that open LLMs can run cost‑effectively on consumer‑grade server hardware with competitive inference speeds.

📬 Get the top 10 AI stories daily