Agent Frameworks

Team of AIs Beats the Best Solo Model in New Test

One AI crew solved 8 of 10 hard math problems — the best solo AI got 6.

Deep Dive

Most AI products you use are a single model doing everything. This new research treats a group of separately running AI models as one service. The authors call the individual models "Cells," and split them into two jobs: Analysts, who can only offer opinions and evidence, and one Executor, who alone is allowed to give the final answer and actually use tools like clicking, typing, or running code. A versioned roster controls who is on the team, so models can be swapped in or pulled out without breaking anything for the app using them.

The results are the interesting part. On a fixed set of hard competition math problems, the AI team solved 8 out of 10. The best single model in the group only solved 6. A saved record showed that knowledge from a minority member — a model most of the team disagreed with — corrected several models that were originally wrong or blank. That is the promise of a team: a quiet teammate catching a mistake everyone else missed.

On 20 real computer terminal tasks, every single tool action traced back to the one Executor. There were zero actions taken by Analysts and zero rules bypassed. During the run, six models were promoted and one incompatible candidate was rolled back while the service kept running — no downtime, no visible change for users. That kind of clean record-keeping matters when something goes wrong and you need to know exactly which AI did what.

The catch: this is a nine-page research paper, not a product you can sign up for. Running several models at once costs more computing power and money than running one, and the strict "one AI does all the acting" rule is a design choice the authors made, not a proven law. Still, it points at where AI is heading — away from one giant brain and toward a managed team with clear rules about who is allowed to do what.

Key Points
  • A research system called Fusion-MoA runs several AI models as one team behind a single plug-in, so existing apps don't need changes.
  • The team solved 8 of 10 hard math problems versus 6 of 10 for the strongest single model — and one model was fixed by a teammate most others disagreed with.
  • Only one AI was permitted to actually take actions, and every action was traceable to it, with six models swapped in and one rolled back without any downtime.

Why It Matters

Signals AI that gives better answers with a clear audit trail of which model did what — crucial when mistakes matter.

📬 Get the top 10 AI stories daily