Open Source

Full Kimi K3 model runs on 16 NVIDIA GB10 cluster at 20+ tps

First community-run full Kimi K3 hits 38 tps peak on 16 GB10 nodes

Deep Dive

Moonshot AI's Kimi K3, a powerful frontier-class language model, has been fully deployed on commodity hardware. In a viral Reddit post, user ciprianveg showcased the model running across a 16-node NVIDIA GB10 cluster (Grace Blackwell Superchip), achieving an average throughput of 20+ tokens per second on a coherent corpus via llama-bench, with peaks at 38 tps and a prefill rate of 750 tps. This marks the first public run of the full Kimi K3 with DSPark, a distributed serving framework, on such a cluster.

The achievement is significant because running a full-size frontier model typically requires expensive data center GPUs. Using a cluster of GB10s—compact, power-efficient modules designed for edge AI—shows that serious inference workloads can be handled by distributed, lower-cost hardware. The author is now optimizing tensor parallelism (TP) to push speeds further and will publish a vLLM image plus setup instructions once ready. This could let researchers and enthusiasts run state-of-the-art models locally without cloud rental fees, while maintaining privacy and full control over the inference stack.

Key Points
  • Full Kimi K3 model served on 16x NVIDIA GB10 cluster via vLLM and DSPark
  • 20+ tps average, 38 tps peak, 750 tps prefill in llama-bench tests
  • Author to release vLLM image and instructions after TP optimization

Why It Matters

Brings frontier-model inference to distributed edge hardware, enabling cheaper, private, and self-hosted AI deployments.

📬 Get the top 10 AI stories daily