Frontier LLMs beat Nash equilibrium in no-communication game tests
Two hosted models exceed game-theoretic baselines without talking, but fail in larger teams
Thirteen language models were tested on one-shot, no-communication multi-agent games spanning seven archetypes and two to ten actions per player. Two frontier-hosted models consistently beat their Nash equilibrium benchmark, approaching optimal joint outcomes in several archetypes, while most open-weight models saw only partial gains that varied sharply by game structure. Performance dropped substantially in team-based games with four or more interchangeable agents, especially as the action space grew, suggesting self-play coordination gains don't transfer beyond dyadic settings.
- Benchmark tested 13 LLMs across 7 archetypes with 2-10 actions per player
- Only 2 frontier-hosted models beat the Nash equilibrium baseline consistently
- Performance dropped sharply in 4+ agent team games, limiting scalability
Why It Matters
Shows current LLM coordination is unreliable without communication, especially in multi-agent systems