Agent Frameworks

Biased consensus emerges in multi-agent LLM debates, ICML study finds

LLM sampling temperature drives debate agents to converge on biased collective norms

Deep Dive

Multi-agent LLM debates are increasingly used for decision-making and problem-solving benchmarks, but their interactive dynamics can amplify inherent biases of individual models. In a paper accepted at the 43rd International Conference on Machine Learning (ICML 2026), researcher Maya Okawa demonstrates that when LLMs debate each other, they tend to form collective norms that are frequently biased. The study identifies noise—especially the LLM sampling temperature—as a key driver of this phenomenon. Building a framework inspired by physics-based models of social dynamics, Okawa predicts a phase transition: once conformity pressure exceeds a critical threshold relative to the agents' initial bias and debate noise, the system flips into a state of collective bias.

Controlled experiments validate this prediction, showing a finite-size crossover consistent with an underlying phase transition. Interestingly, agent heterogeneity—having models with diverse biases or behaviors—smooths the transition and suppresses the emergence of biased consensus. The insights extend beyond toy settings: tests on realistic decision-making tasks, including investment choices and LLM-as-a-judge evaluation, confirm the same dynamics. This work highlights a substantial safety and fairness risk for multi-agent LLM deployments, where otherwise balanced models can be nudged into extreme or skewed positions simply by the debate process. For practitioners, it implies that careful tuning of debate parameters, such as temperature and participant diversity, is essential to avoid biased outcomes in real-world applications.

Key Points
  • Noise (e.g., LLM sampling temperature) is a key driver of collective bias in multi-agent debates
  • Physics-inspired model predicts a phase transition when conformity exceeds a critical bias threshold
  • Agent heterogeneity suppresses bias; findings apply to investment decisions and LLM-as-a-judge

Why It Matters

Multi-agent LLM deployments risk biased decisions; understanding phase transitions helps design safer, fairer debate systems.

📬 Get the top 10 AI stories daily