New Study: AI's Smarts Can't Be Measured by One Score
AI aces one task and flops the next — new research explains why.
Three researchers analyzed 13,251 published test scores covering 1,618 AI language models and 456 different exams. They borrowed a method psychologists use to study human intelligence — looking for a hidden 'general factor' (a single underlying smartness) behind all those scores. Their conclusion: such a factor exists, but it explains a lot less than AI companies assume. At their most generous estimate it accounted for about 71% of performance, and far less in most of their calculations.
Two other findings are just as interesting. First, tests that look similar on the surface don't necessarily measure the same skill — passing one doesn't reliably predict passing another. Second, the general factor isn't dominated by any one theme, and the tests commonly used as shorthand for 'intelligence' turn out to be poor stand-ins for it. In plain terms: there's no single dial labeled 'smart' that you can turn up.
That matters because a lot of AI development today is chasing exactly that dial — trying to build one model that gets broadly better at everything. If abilities are partly idiosyncratic and don't cluster neatly, targeting a single overall capability doesn't work well. Progress will likely stay uneven: a model that improves at summarizing contracts may not improve at reading a chart or handling customer complaints.
So what should you do with this? Treat benchmark scores in AI marketing the way you'd treat a restaurant's star rating from a critic who never ate your favorite dish. Before trusting an AI tool with real work, test it on your actual task, with your actual documents. Expect surprises in both directions — moments of brilliance and baffling mistakes from the same tool. And when someone promises that one model is about to become universally smarter, remember: this study suggests there's no single score for that.
- One overall 'smartness' score explains only about 70% of how AI models perform — the rest is task-specific.
- The study looked at 13,251 scores from 1,618 AI models on 456 different tests.
- Tests that look similar don't reliably measure the same skill, so a high score on one may not predict another.
Why It Matters
Expect AI to stay uneven — test it on your own task, and treat 'smarter AI' marketing claims skeptically.