New Math Shows Exactly How Much Accuracy AI Loses When You Train It on Less Data
A precise recipe for training AI with fewer examples — and less cost.
Training an AI usually means feeding it examples: emails to spot spam, houses to predict prices, clicks to guess what you'll buy next. A natural question is whether you really need all of them. Could you use half the data — picked cleverly — and get almost the same quality? Until now, the honest answer was "probably, but nobody can tell you exactly how much worse it gets." This new paper supplies that exact answer for a common, well-understood type of model.
The authors prove a simple rule. If your model looks at a certain number of separate inputs (think of these as dials, like age, income and past purchases), and you train on a smaller set of examples, the extra error follows a tidy formula: it shrinks predictably as you add more examples. They also show their formula is the best possible — you cannot do better with a cleverer selection method, and they include examples proving it. Because the result is a formula rather than a rough estimate, engineers can plan ahead instead of guessing.
What makes this unusual is how it was checked. The authors ran the entire proof through a program called Lean 4, which verifies mathematical logic the way a spell-checker verifies spelling — it either confirms every step or refuses. That matters because AI research has a quiet credibility problem: results can be hard to reproduce, and small errors in math can hide for years. A machine-checked proof is about as trustworthy as mathematics gets.
So should you expect your apps to get better next week? No. This is plumbing, not a product. It describes a specific family of simple models, not the giant language models behind chatbots. But plumbing is what makes the rest work: every dollar and hour spent training AI comes down to how much data is enough. Knowing the exact answer turns an expensive argument into a calculation — and cheaper training eventually shows up as cheaper tools for everyone.
- Researchers found an exact formula for how much accuracy a model loses when trained on a smaller, carefully chosen slice of data — no more educated guessing.
- The formula was verified line by line by Lean 4, a computer program that checks math proofs, making the result unusually trustworthy for AI research.
- It applies to simple prediction models, not today's big chatbots, so expect cheaper and faster training pipelines before any visible product change.
Why It Matters
Knowing exactly when less data is enough could cut AI training costs — eventually making AI tools cheaper and faster for everyone.