LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability
Chenxu Wang, Yongkun Yang, Boyuan Du, Shiwei Lin, Huaping Liu
Why It Matters
What makes this one worth your time
Understanding the capabilities and limitations of LLMs in collaborative settings is crucial for developing more effective multi-agent systems and improving AI-human interactions.
The paper evaluates LLM agents in collaborative decision-making tasks, highlighting their challenges and potential for improvement.
Summary
The paper investigates the performance of large language model (LLM) agents in deliberative collaboration tasks under partial observability. It formalizes the problem as a cooperative joint decision-making task and introduces a scalable benchmark to evaluate LLMs across various domains. The study finds that while current LLMs struggle with complex deliberative tasks, the deliberation process can sometimes enhance performance through reflection and error correction.
Key contributions
- Formalization of deliberative collaboration as a joint decision-making problem with partial observations.
- Introduction of a scalable benchmark for evaluating LLMs in deliberative tasks.
- Systematic evaluation of representative LLMs in collaborative settings.
Notable insights
- Deliberation can serve as a mechanism for reflection and error correction, potentially improving performance over centralized baselines.
- Complex deliberative tasks remain challenging for state-of-the-art LLMs, even with external mathematical tools.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2607.06157v1 Announce Type: cross Abstract: Deliberation plays a crucial role in collaboration; when humans work together, they naturally engage in communication to align information and reach an agreement. In this paper, we investigate deliberative large language model (LLM) agents under partially observable joint decision-making tasks. We formalize deliberative collaboration as a cooperative joint decision problem with partial and asymmetric observations, and introduce a scalable benchmark that instantiates this problem across multiple task settings and domains in which agents must exchange information through deliberation to reach a joint decision with a shared reward. We then instantiate a reference scaffold and evaluation protocol for deliberative agents and conduct a systematic evaluation of a range of representative LLMs. The results reveal that complex deliberative collaboration tasks continue to challenge state-of-the-art language models. Even with the aid of external mathematical tools, language models may fail in either the deliberation process for aligning information or the complex reasoning process for making the decision. On the other hand, diagnostic analysis reveals that the deliberation process may also provide opportunities for reflection and error correction, sometimes improving performance over centralized baselines. Altogether, our work establishes a foundation for evaluating and improving LLM agents in deliberative collaboration and provides insights into the strengths, limitations, and properties of current LLM-based multi-agent systems.