Open Source

MiniMax 2.7 agent model hits 47 tg/s on 96GB VRAM rig

A Reddit user runs a powerful agent-class LLM locally with multi-agent loops and 192GB RAM.

Deep Dive

Redditor Important_Quote_1180 runs the MiniMax 2.7 REAP Q4 on 96GB VRAM, 192GB DDR5 UDIMM, a B840 MSI board, a 9900X CPU, and Ubuntu Linux with a 1250W PSU (power limited). The model is an agent-class model with excellent instruction following and tool calling. They use a round-robin loop of 3 sequencing agents running on the CPU (MoE models at 15–20 tg and 300 PP each) with 20–40k token system prompts, plus an asynchronous dense 12B model that watches the loop and flags one issue. Each loop takes 4–10 minutes.

Key Points
  • MiniMax 2.7 runs at 47 tg/s on 96GB VRAM with REAP Q4 quantization
  • Uses 3 CPU-based MoE sequencing agents in a round-robin loop (15–20 tg, 300 pp)
  • Includes a 12B dense model for asynchronous monitoring of the entire pipeline

Why It Matters

Shows it's possible to run advanced agent-class AI locally with multi-agent coordination and high throughput.

📬 Get the top 10 AI stories daily