Open Source

Xiaomi's MiMo V2.5 hits 3000 tps with DFlash & persistent kernel

Open-source release promised after hitting 3000 tokens/sec inference.

Deep Dive

Xiaomi has unveiled MiMo V2.5, a new model architecture that pushes inference speeds to an unprecedented 1000-3000 tokens per second (tps). The key innovations are DFlash (direct flash attention) and a persistent kernel design that minimizes memory overhead and maximizes GPU utilization. Unlike typical large language models that trade off speed for quality, MiMo V2.5 maintains competitive benchmarks while achieving sub-millisecond per-token latencies.

Already, Xiaomi has released the model weights under the MIT license on their blog, and they promise a full open-source release including training code and evaluation scripts soon. This development could democratize real-time AI applications—chatbots, code assistants, and interactive agents—by slashing inference costs and latency. For professionals, this means running state-of-the-art models on affordable hardware without cloud dependencies is rapidly becoming a reality.

Key Points
  • MiMo V2.5 achieves 1000–3000 tokens per second using DFlash and persistent kernel
  • Model weights are publicly available; full open-source release (training code) promised soon
  • Enables real-time inference on consumer GPUs, reducing cloud dependency for LLM applications

Why It Matters

Sub-1ms per token opens real-time AI on local hardware, challenging cloud dominance.

📬 Get the top 10 AI stories daily