Open Source

WASTE runs 2.78T-parameter Kimi K3 from NVMe without GPU memory

Streams activated weights from NVMe, so 2.78T parameters fit in any rig

Deep Dive

WASTE is an embeddable inference engine written in C, with no third-party runtime dependencies. It keeps the model trunk in memory, streams selected experts directly from disk, and uses the remaining RAM as a bounded expert cache.

Key Points
  • Streams only activated MoE experts from NVMe, slashing RAM requirements for 2.78T-parameter Kimi K3
  • Written in dependency-free C, making it embeddable in any application or edge device
  • Bounded expert cache uses remaining RAM to keep hot experts fast, balancing I/O and memory

Why It Matters

Makes trillion-parameter MoE models accessible on modest hardware, democratizing frontier AI inference.

📬 Get the top 10 AI stories daily