New Trick Cuts Giant AI Models in Half Without Losing Smarts
Half the memory, 93% of the brainpower — cheaper, faster AI on the way.
Big AI models like DeepSeek and Mixtral are built like a giant hospital full of specialists. Instead of one doctor answering everything, hundreds of tiny expert brains sit on standby, and only a few get picked for each question. It's a clever design — but keeping all those specialists on the payroll costs a fortune in memory and computing power, which is why running top AI is so expensive.
The new method, called OMP-MoE, is a no-retraining cleanup crew. It watches which experts actually get used and which just sit around collecting a salary, then fires the redundant ones. The authors use a math shortcut (borrowed from signal processing) to find the best small team of experts quickly, then balance how many experts each layer of the model keeps. A bonus feature adjusts on the fly, activating experts only when they're needed.
The results are striking. On Qwen3-30B-A3B, cutting the expert count in half kept 93.3% of the original performance — and searching for which experts to cut became 33 times faster. Answers came back 1.55x quicker. They tested it on Qwen, DeepSeek-V2, GPT-OSS and Mixtral, and it beat existing trimming methods consistently at both 25% and 50% cuts.
So what? Cheaper hosting usually means cheaper AI subscriptions, faster chatbots, and models that can run on smaller machines — potentially your laptop or phone instead of a data center. The catch: this is a preprint (not yet peer-reviewed), and the code won't be released until it's accepted. It's promising, but not yet something you can use today.
- Mixture-of-experts AI stores many specialist sub-brains and uses a few per question — powerful, but memory-hungry and costly to run.
- OMP-MoE trims half of those specialists without retraining, keeping 93.3% of performance on a 30-billion-parameter model.
- It made the model respond 1.55x faster and sped up the search for which experts to cut by 33 times — pointing to cheaper, snappier AI services.
Why It Matters
Cheaper, faster AI models could lower subscription prices and let powerful assistants run on everyday laptops and phones.