Agent Frameworks

LLM agents fail at consensus without classical filters, study finds

New research: prompted LLM agents can't agree even where theory guarantees convergence

Deep Dive

Researchers at the University of Pennsylvania, Sribalaji C. Anand and George J. Pappas, have published a study on arXiv asking whether classical resilient consensus theory — developed for deterministic agents — transfers to LLM agents that may behave adversarially. Framing LLM agreement as a Byzantine consensus game, they ran controlled experiments on complete and general communication graphs.

Their key finding: prompted LLM agents fail to reach agreement that is achievable in principle. Consensus can fail even in settings where classical theory guarantees that a convergent algorithm exists, and this failure persists across different temperatures and time horizons. However, wrapping the agents with classical resilient consensus filters improves agreement. The benefit of filtering depends on how much robustness the underlying topology already provides. The results suggest classical resilient consensus theory is a useful lens for the safety of agentic AI.

Key Points
  • Prompted LLM agents fail to achieve consensus in settings where classical theory guarantees convergence.
  • Failure persists across temperatures and horizons, showing the issue is fundamental, not just parameter-dependent.
  • Wrapping agents with classical resilient consensus filters improves agreement, but benefit depends on graph topology robustness.

Why It Matters

Classical consensus theory can enhance safety and reliability of multi-agent LLM systems.

📬 Get the top 10 AI stories daily