Open Source

Xiaomi's MiMo-V2.5-Pro UltraSpeed claims 1,000+ tps on 1T MoE model with 8 GPUs

Xiaomi says it broke 1,000 tokens/sec on a 1-trillion parameter model using just 8 GPUs…

Deep Dive

Xiaomi MiMo, the AI division of the Chinese electronics giant, has announced MiMo-V2.5-Pro UltraSpeed, claiming a staggering throughput of over 1,000 tokens per second on a 1-trillion parameter mixture-of-experts (MoE) model. The benchmark is notable because it was achieved on a single standard 8-GPU node, without resorting to custom wafer-scale chips (like Cerebras) or SRAM-laden architectures (like Groq). The company asserts this breaks the elusive 1,000 tokens/sec barrier for models of this scale, a milestone previously thought to require exotic hardware.

If the claims hold up under independent verification, the implications for enterprise AI infrastructure are profound. Running massive MoE models on commodity GPU servers would slash deployment costs and power consumption, making trillion-parameter inference accessible to far more organizations. However, skepticism is warranted—Xiaomi has not released detailed benchmarks or methodology. The AI community will be watching closely for replication attempts or third-party evaluations.

Key Points
  • Xiaomi MiMo claims 1,000+ tokens/sec throughput on a 1-trillion parameter MoE model.
  • Achieved using a single standard 8-GPU node, not custom hardware like Cerebras or Groq.
  • If validated, this would dramatically lower hardware and cost barriers for large model inference.

Why It Matters

Could democratize trillion-parameter model inference by removing the need for exotic, expensive hardware.

📬 Get the top 10 AI stories daily