New Test Shows AI Teams Fail Half the Time at Working Together
Even the best AI can't team up: they fail nearly half the time.
A group of university researchers just released AgentWorld, an open-source test that puts AI teams inside a video-game world — think World of Warcraft, not a spreadsheet. Each test has 100 hand-written tasks, and teams of 3 to 20 AI programs (software that can take actions on its own) must work together for 50 or more rounds to finish them. The twist: each AI acts alone and cannot see what its teammates are thinking.
The results are humbling. Four leading models took the test — Google's Gemini 3 Flash, Anthropic's Claude Haiku 4.5, OpenAI's GPT-5 Mini, and China's DeepSeek R1-70B. The best of them completed only 52% of the tasks. That's barely better than a coin flip. The failures followed clear patterns: teams stopped talking to each other, forgot the plan halfway through, or got confused about who was supposed to do what.
The researchers also invented a smarter way to grade teamwork. Instead of just asking "did they win, yes or no?", their new measure traces which actions actually mattered — like checking how much of a group project's work really made it into the final report. This exposes wasted effort that a simple pass-or-fail score hides.
Why should you care? AI "agents" are being marketed right now as digital coworkers that can handle multi-step chores: booking travel, sorting email, managing orders. This test shows the teamwork is still shaky. The practical takeaway: keep a human in the loop on anything long or complicated, and don't expect a team of AI assistants to run a project unsupervised. The good news is the test is free and open, so progress can now be measured instead of guessed.
- The test drops teams of 3 to 20 AI programs into a video-game world where they must cooperate for 50+ rounds — like a group project where nobody can see each other's notes.
- The best model scored just 52% — Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini and DeepSeek R1-70B all stumbled on communication and forgotten plans.
- Researchers created a new score that measures how much of a team's effort actually contributed to the final goal, revealing hidden wasted work.
Why It Matters
Before you let AI assistants run your inbox, bookings or projects unsupervised, know they still fail half the time at teamwork.