Back to today's list

When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains

Chen Liang, Fasheng Xu

Published Aug 11, 2026Featured #9In the daily list Aug 12, 2026
Daily score66.2
Editorial review7.2
Relevance0.505
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding how LLMs perform in autonomous negotiations can guide firms in deploying these models for procurement, ensuring efficient and profitable outcomes.

The study benchmarks LLMs in supply chain negotiations, revealing insights into value creation and surplus division.

Summary

The paper evaluates the performance of nine large language models (LLMs) from OpenAI, Google, and Alibaba in a supply chain bargaining scenario, focusing on value creation, surplus division, and contract reliability. It highlights the role of capability, provider identity, and strategic prompt design in negotiation outcomes.

Key contributions

  • Benchmarking LLMs against a Perfect Bayesian Equilibrium in supply chain negotiations.
  • Identifying the impact of provider identity and strategic prompt design on negotiation outcomes.

Notable insights

  • Provider identity influences surplus capture more than capability rank, suggesting vendor choice is crucial.
  • Strategic prompt design significantly impacts surplus division, indicating a key area for optimization.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2608.07538v1 Announce Type: new Abstract: As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictably, and avoid money-losing contracts. We study this in a canonical supply chain bargaining problem: a buyer with private demand information negotiates a quantity-payment contract with an uninformed seller. We benchmark nine LLMs from OpenAI, Google, and Alibaba against a validated Perfect Bayesian Equilibrium across 9,840 LLM-to-LLM negotiations. First, capability governs value creation. Agents agree in 98.9% of negotiations and capture 95.4% of first-best surplus undiscounted, but average 2.98 rounds against the benchmark's 1.25, and this delay erodes 21-34% of surplus. Capability also governs reliability: baseline models accept individually irrational contracts in 19.2% of cases, versus 0.0-0.6% at mid-tier and flagship, making automated profit verification the binding guardrail below that threshold. Second, surplus capture is relational. Provider identity predicts who captures surplus better than capability rank: self-play buyer shares average 40% for OpenAI, 50% for Google, and 70% for Alibaba's Qwen, an ordering that survives restricted communication and no discounting. Reversing which provider sells moves the division by 7-18 percentage points, and the capable Qwen flagship is the weakest cross-family seller: vendor choice is a first-order distributional decision. Third, the prompt is a strategic lever. Delegation separates the principal's economic patience from the agent's prompted strategic patience, a free deployment choice that is the single strongest driver of surplus division (90% of explained variance). Together these establish an equilibrium-referenced audit of AI agents along three dimensions: discounted efficiency, distributional profile, and operational reliability.