Scientists Finally Explain Why AI Trains So Well Even When the Math Is Wrong
The engine behind ChatGPT takes mathematical shortcuts — and still wins.
Every time you hear that a new AI model was "trained," a piece of software called Adam was doing the steering. Adam is the default driving instructor for almost all modern AI, from image generators to chatbots. Think of it as a driver trying to reach a destination. There's a theoretically perfect route, called natural gradient descent. Adam doesn't take that route. It takes shortcuts, and this paper measures how big those shortcuts are.
The researcher tested Adam across four problems, from simple to messy, including a small neural network. On easy, well-behaved problems, Adam stayed fairly close to the ideal route. On messy problems, it drifted badly — at one point misaligned by roughly a thousand times. That sounds alarming, but here's the twist: the drift slowed Adam down at the start without hurting the final result. Adam still finished with an excellent answer every time.
The second surprise was about momentum — Adam's habit of 'keeping rolling' in a direction rather than turning sharply. That smoothing turns out to be a big part of why Adam works. The study also found that a refined version of its internal guesswork follows a steadier path than the standard version, which often wobbles or blows up entirely.
So what does this mean for you? Nothing changes in your apps tomorrow. But training AI is enormously expensive, and that cost eventually shows up in subscription prices and energy bills. Understanding why Adam succeeds despite imperfect math could help researchers design faster, cheaper training. It also explains why AI is surprisingly tough: it doesn't need perfect math to be useful, just math that's good enough, applied consistently.
- Adam is the standard math recipe that trains almost all modern AI, including chatbots and image tools
- On messy problems it drifts far from the 'perfect' path — up to about a thousand times off — but still reaches good answers
- Its real strength seems to be momentum (smoothing out bumps) rather than mathematical precision, which hints at cheaper, faster AI training ahead
Why It Matters
Could lead to cheaper, faster AI training — savings that eventually reach your subscriptions and energy bills.