New Software Makes Giant AI Twice as Fast on the Same Chips
Faster chatbot replies and cheaper AI subscriptions — without buying any new hardware.
Big AI models like DeepSeek don't use their whole brain for every question. They use a trick called mixture-of-experts (MoE) — think of a huge company where only the relevant departments answer each request. That saves power, but it creates a new problem: those departments live on different computer chips, and some chips get flooded with work while others sit around doing nothing. The whole system can only move as fast as its busiest chip, so all that idle time is wasted money.
The MegaFlux team, from Stanford and NVIDIA among others, fixed this by letting the system decide at runtime which chips should take on extra work. Popular "experts" get copied onto idle chips on the fly, and the copies start working the moment they receive the needed data instead of waiting for everything to arrive. They tested it on eight NVIDIA B200 chips across 147 different setups: about 1.45 times faster on the main processing step and 1.28 times faster on the training step, with peaks of 2.14 and 2.64 times. In a real chatbot system running DeepSeek-V4-Pro, users saw roughly 13–26% faster responses end to end.
Why should you care? Every second an AI chip sits idle is a second you're paying for. Speeding up existing hardware by 20–50% is like getting a bigger data center for free — it lowers the cost of running chatbots, coding assistants, and search. Those savings tend to show up as cheaper subscriptions, faster replies, and the ability to serve more people at once during peak hours.
The catch: this is research code tested on top-tier, expensive chips, and it only helps models built in this "many experts" style. Smaller AI models won't benefit, and companies still have to adopt the software before you notice anything. But it points to a trend worth watching — the next big AI cost savings may come from clever plumbing, not bigger chips.
- MegaFlux stops expensive AI chips from sitting idle by moving work to whichever chip is free at that moment.
- Tested on eight NVIDIA B200 chips, it made AI processing 1.45x faster on average and up to 2.14x faster at best.
- In a real chatbot setup with DeepSeek-V4-Pro, responses got 13–26% faster without any new hardware.
Why It Matters
Cheaper, faster AI means lower subscription prices and quicker answers, since providers can serve more users per chip.