Research & Papers

New Method Trains Big AI Twice as Fast — Cheaper AI Could Follow

⚡Faster, cheaper AI training usually means lower prices and quicker features for you.

Deep Dive

Big AI models today are not one giant brain. They are built like a huge company of specialists — one for math, one for cooking questions, one for coding — and for any given question, only a few specialists wake up and work. That design is called "mixture of experts" (many small expert models, only a few active at once). It lets companies build enormous AI without enormous electricity bills.

The catch is that those specialists live on different computer chips, sometimes in different buildings. Every time you ask the AI something, data has to be shipped back and forth between the chips. That shipping is slow, and it's one of the main reasons training a top AI model can cost tens of millions of dollars.

The researchers noticed something simple: certain specialists almost always get used together. So they built Cobalt, which plays matchmaker — it moves the specialists that work as a team onto the same chip, and re-sorts them as the training goes on. Testing on 32 of Nvidia's newest B200 chips, Cobalt trained models 1.5 to 2.4 times faster than standard methods and cut data shipping between machines by 75% to 99%.

So what? Training speed is the main cost driver behind AI. When a company can train in half the time, it can release better models sooner, charge less, or run more experiments. You rarely see a paper like this directly, but you feel it: the AI assistant you use next year will likely be trained with tricks like this. The honest limitation is that Cobalt was tested at research scale on one specific type of expensive hardware, and real-world gains at massive scale are still unproven. Still, the direction is clear — the plumbing behind AI is getting much more efficient.

Key Points
  • Cobalt is a technique that trains big AI models up to 2.4 times faster by keeping frequently paired 'specialist' parts on the same chip.
  • It cut data shuttling between machines by up to 99% in tests on 32 Nvidia B200 chips, the newest AI hardware.
  • The biggest AI models are split across many chips, so this kind of plumbing work directly lowers the cost of building AI.

Why It Matters

Cheaper, faster AI training often means lower prices and quicker improvements in the AI tools you use daily.

📬 Get the top 10 AI stories daily