Research & Papers

MuZero dominates asymmetric Baghchal with 86% tiger win rate

Deep RL beats Nepal's strategic game of 4 tigers vs. 20 goats

Deep Dive

A new paper from Ranjit Raut and colleagues systematically applies four deep reinforcement learning algorithms—DQN, REINFORCE, PPO, and MuZero—to the asymmetric board game Baghchal, which pits four tigers against twenty goats. The game, of Nepali origin, has rich strategic depth and perfect information but has been understudied in RL literature. Each algorithm was trained on one side of the asymmetry and evaluated on both, with metrics including win rate, draw rate, average captures, training convergence, and computational cost.

MuZero emerged as the clear winner, achieving 86% wins as tigers and 62% as goats, leveraging its model-based Monte Carlo Tree Search for long-horizon planning. PPO proved a practical, hardware-friendly alternative with competitive performance at much lower compute cost. Notably, DQN exhibited a strong bias toward the tiger role, attributed to the larger reward signal for tigers. The findings highlight that model-based planning excels in asymmetric games, offering insights for broader AI game-playing and strategic decision-making.

Key Points
  • MuZero achieved 86% win rate as tigers and 62% as goats, the highest across all algorithms
  • PPO provided near-competitive performance with significantly lower computational cost than MuZero
  • DQN showed bias toward the tiger role due to stronger reward signals, limiting goat-side play

Why It Matters

Demonstrates that model-based RL (MuZero) excels in asymmetric strategy, with implications for real-world adversarial planning and game AI.

📬 Get the top 10 AI stories daily