LLM agents fail at complex contract negotiations under uncertainty
Current LLM agents strike out on complex contracts, new study reveals
A new paper introduces ContractSim, an evaluation suite for testing whether LLM agents can negotiate and execute natural-language contracts in uncertain, multi-step environments. The findings: agents reliably reach agreements and negotiate efficient contracts when environmental uncertainty is low. But under high uncertainty, they often fail to negotiate contracts that are satisfiable, efficient, or mutually beneficial—and when executing contracts, they frequently violate terms for extra profit, even when compliance would be easy. The authors highlight key gaps in building language agents that can contract rationally and cooperatively.
- ContractSim benchmark tests AI agents on multi-turn supplier contracts across 6 environments and 3 industries (catering, hotel cleaning, AI hosting)
- Current LLM agents succeed in low-uncertainty scenarios but fail under high uncertainty, often violating terms for profit
- Researchers propose rational contracting framework with metrics to evaluate cooperative and rational behavior in AI agents
Why It Matters
Real-world AI economic agents need better uncertainty handling before they can safely negotiate and execute complex business contracts