AI 'Expert Panels' Can Agree and Still Be Wrong, Study Finds
A room full of AI agreeing isn't proof it's right — here's why that matters.
Companies increasingly use teams of AI chatbots — often called multi-agent systems, meaning several AIs working together — to deliberate, review work, or act as a judging panel. The common assumption is simple: the more the AIs talk to each other, the more reliable their shared answer becomes. This study tested that assumption by running 50 AI agents that each updated their answer based only on a few neighbours they could "see," then watched how agreement formed over 20 rounds.
The team found three patterns. Sometimes everyone lines up together (they call it synchronised). Sometimes the panel looks tidy in small clusters but the overall group is incoherent — a "twisted" state, like a rumour that sounds consistent in every neighbourhood but contradicts itself across town. And sometimes you get a chimera: half the panel agreeing, half disagreeing, side by side. Worryingly, 40% of trials that looked extremely consistent still ended in a twisted state at the end.
The two dials do different jobs. Turning up the AI's thinking effort (giving the model more time to reason) pushed panels away from messy, fragmented answers toward those locally orderly but globally inconsistent states. Turning up how many peers each AI could see pushed panels toward genuine group-wide agreement. That second effect held across three different AI companies' models and on a separate judging task.
The catch: this is a simulation, not a real-world audit, and it tells us how agreement forms — not who is actually correct. A panel can be perfectly aligned and completely wrong. If your workplace treats AI consensus as a quality check, ask how the panel was wired before trusting the verdict.
- Fifty AI chatbots were wired together and watched as they changed their answers — agreement often formed in misleading patterns rather than genuine consensus.
- Making the AI think harder didn't make it more accurate; it just made the panel's mistakes look more organised and consistent.
- Groups that looked extremely unified (about 40% of trials) still ended up internally contradictory, so 'everyone agrees' is a weak safety signal.
Why It Matters
If your team trusts AI panels to double-check decisions, shared agreement may be hiding a wrong answer.