More AI Agents Doesn't Mean Better Answers, Study Finds
Hiring 30 AI helpers can cost 30x more — but rarely works 30x better.
Companies building AI products love the idea of "agents" — AI helpers that work in teams, checking each other's work and splitting up jobs. The logic seems obvious: more helpers, more brainpower. Two researchers in Slovenia tested that assumption by running 13 downloadable AI models in teams of up to 30 copies each, across a range of test questions. Their conclusion: it depends entirely on what kind of job the team is doing.
On what they call "needle" tasks — where you just need one member to get it right, like finding a rare error or generating a clever idea — bigger teams genuinely worked. The odds that at least one AI got it right rose by 5 to 20 points as the team grew. But here's the catch: once you try to pick the winning answer by majority vote, you get almost none of that benefit. The group just defaults to whatever a single AI would have said anyway, coming within half a point of it on average. More voices, same answer.
The one technique that clearly helped was letting the AIs revise their answers after seeing a peer's response. That produced real accuracy gains — but a single teammate was nearly as good as 29 of them. On rough estimation questions ("how many piano tuners are in Chicago?"), teams were nearly useless. Averaging answers cut the error by only about 6%, because each AI carries its own built-in bias, which accounted for roughly 87% of the total error. Adding more copies of a biased AI just repeats the same bias.
Mixing different AI brands helped on estimation questions, but still didn't beat the single best model on the needle-type tasks. The practical takeaway: bigger AI teams aren't automatically smarter. What matters is matching the team structure — and how you combine their answers — to the job you're actually doing.
- More AI helpers only helps on tasks where one correct answer is enough — and even then, letting them vote cancels out most of the gain.
- Letting AI agents revise after reading a teammate's answer improves accuracy, but one teammate works nearly as well as 29.
- Stacking 30 copies of the same AI wastes money on guess-the-number tasks: shared bias caused about 87% of the error.
Why It Matters
If your company is paying for armies of AI agents, this suggests fewer, better-organized ones often do the job.