Research & Papers

LLM agents nail 95.4% of deal value but waste 21-34% on slow talks

9,840 AI-to-AI negotiations reveal Qwen grabs 70% of surplus vs OpenAI's 40%.

Deep Dive

A new arXiv paper from Chen Liang and Fasheng Xu puts nine LLM agents from OpenAI, Google, and Alibaba through 9,840 head-to-head procurement negotiations, using a canonical supply chain bargaining game where a buyer with private demand information haggles with an uninformed seller. Against a validated Perfect Bayesian Equilibrium benchmark, the agents reach agreements 98.9% of the time and capture 95.4% of the theoretical first-best surplus undiscounted. But efficiency drops sharply in practice: negotiations average 2.98 rounds versus the benchmark's 1.25, and that delay destroys 21-34% of surplus. More concerning, baseline models accept individually irrational contracts in 19.2% of cases, while mid-tier and flagship models fail only 0-0.6% of the time, suggesting automated profit verification is essential below a certain capability threshold.

The paper's most surprising finding is that surplus capture is less about raw capability and more about provider identity. In self-play, OpenAI models keep 40% of the surplus, Google models 50%, and Alibaba's Qwen models 70%—an ordering that persists even under restricted communication and no discounting. Switching which provider supplies the buyer or seller shifts profit division by 7-18 percentage points, and the capable Qwen flagship is the weakest cross-family seller, meaning vendor selection is a first-order distributional decision. Prompt engineering also turns out to be a strategic lever: by separating the principal's economic patience from the agent's prompted strategic patience, companies can shift surplus division dramatically—the prompt alone explains 90% of the variance in outcomes. The authors frame this as an equilibrium-referenced audit of AI agents across discounted efficiency, distributional profile, and operational reliability.

Key Points
  • 98.9% deal agreement rate but 2.98 rounds vs optimal 1.25, eroding 21-34% of surplus
  • Baseline LLMs accept money-losing contracts 19.2% of the time; flagships drop to 0-0.6%
  • Provider identity beats capability: Qwen self-play captures 70% surplus vs OpenAI's 40%, Google's 50%
  • Prompt design drives surplus division, explaining 90% of variance across negotiations

Why It Matters

As AI takes over procurement, choosing the right model provider and prompt can shift billions in value—and unguarded agents will lose money.

📬 Get the top 10 AI stories daily