Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game
Niklas Bauer, Lars Benedikt Kaesberg, Akiko Aizawa, Jan Philip Wahle, Bela Gipp, Terry Ruas
Why It Matters
What makes this one worth your time
Understanding LLMs' capabilities in deception and reasoning is crucial for ensuring their safe deployment in high-stakes environments.
ParliamentBench evaluates LLMs' deceptive and reasoning abilities using a social deduction game framework.
Summary
The paper introduces ParliamentBench, a benchmark framework based on the social deduction game Secret Hitler, to evaluate large language models (LLMs) in scenarios requiring deception, persuasion, and reasoning under information asymmetry. It assesses 16 LLMs across 1,600 simulated matches, introducing three novel metrics to measure social deduction, reasoning, and deceptive consistency.
Key contributions
- Development of the ParliamentBench framework for evaluating LLMs in deception and reasoning tasks.
- Evaluation of 16 LLMs across 1,600 simulated matches with humans and against online games.
- Introduction of three novel metrics for assessing social deduction, reasoning, and deceptive consistency.
Notable insights
- The use of a social deduction game as a reproducible proxy to study adversarial behaviors in LLMs.
- Introduction of novel metrics to isolate and evaluate social deduction and deception consistency in LLMs.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2607.28146v1 Announce Type: new Abstract: As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safety. Controlled social deduction games provide a reproducible proxy for isolating and evaluating these complex adversarial behaviors. We present the open-source benchmark framework ParliamentBench based on the game Secret Hitler to evaluate LLMs in scenarios that require deception, persuasion, and reasoning under information asymmetry. We evaluate 16 LLMs across 1,600 simulated matches playing each other, playing against humans, and compare them against a large set of online games. We introduce three novel metrics that isolate social deduction, reasoning, and deceptive consistency. Our experiments reveal that frontier models achieve strong performance across cooperative and deceptive roles, with a strong top-four cluster (GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus), whereas the weakest models fall short of random (33%) and simple algorithmic (45%) baselines. Most LLMs struggle to maintain a consistent deceptive persona throughout an entire game, with deception retention dropping below 50%.