AlphaZero study reveals limits in optimal play for Connect Four and Chomp
New research shows even superhuman AI can't always play perfectly.
AlphaZero has demonstrated superhuman performance in games like Chess and Go, but new research from Brent Kong, Tejas Ram, and Tony Yue Yu reveals a critical gap: strong play does not equal perfect play. The team tested AlphaZero on Connect Four (a solved partisan game) and Chomp (an impartial game with Grundy-number structure). Using a unified self-play plus MCTS pipeline, they found that vanilla AlphaZero achieves high win rates but fails to preserve exact optimal trajectories. In Connect Four, it cannot maintain the optimal line of play; in Chomp, it fails to consistently restore the g=0 invariant required for perfect play. Even multi-frame inputs on rectangular Chomp boards did not close this gap.
To bridge the gap, the authors introduced AlphaZero Auxiliary Loss (AZAL), which adds oracle-derived policy supervision during training. AZAL substantially improved oracle consistency across multi-seeded full-game traces and sampled-state evaluations. On Chomp, AZAL achieved perfect full-game oracle consistency on 10x11 boards and high (but not complete) consistency on 9x10 boards. On Connect Four, AZAL improved the oracle-match rate and delayed the first oracle mistake, but did not reach perfect play. These results suggest that while auxiliary supervision helps, achieving perfect play in complex games may require additional architectural innovations or more fundamental changes to the training paradigm.
- Vanilla AlphaZero fails to maintain optimal play in Connect Four and Chomp despite strong performance.
- The proposed AZAL method adds oracle-derived policy supervision, improving consistency but not achieving perfect play in Connect Four.
- On 10x11 Chomp boards, AZAL reached perfect full-game oracle consistency; on 9x10, high but not complete.
Why It Matters
Highlights the gap between strong and perfect AI play, crucial for safety-critical applications.