Back to today's list

Auditing Belief-Conditioned LLM Agents in Hidden-Information Social Deduction Games

Yuan Gao, Jiangyi Yang, Yao Zhao, Yichi Zhang

Published Jul 14, 2026
Editorial review6.8
Relevance0.497
Freshness0.000

Why It Matters

What makes this one worth your time

Understanding and improving decision-making in hidden-information games can enhance the development of more robust and interpretable AI agents, which is crucial for applications requiring strategic reasoning and collaboration.

An auditable framework for evaluating belief-conditioned LLM agents in hidden-information games reveals improved outcomes but unresolved mechanisms.

Summary

The paper presents an auditable framework for evaluating belief-conditioned large language model (LLM) agents in hidden-information social deduction games, specifically in a 9-player Werewolf environment. It introduces an external belief state to log belief updates and belief-action deviations, supporting a defensive offline improvement loop. The study finds that active-belief conditions improve good-side outcomes, but the mechanism remains unresolved. The framework helps measure effects, exposes low action-belief consistency, and separates strategy effects from load confounds.

Key contributions

  • Development of an auditable framework for evaluating LLM agents in hidden-information games.
  • Demonstration of improved good-side outcomes under active-belief conditions.
  • Separation of strategy effects from load confounds in the evaluation process.

Notable insights

  • The framework logs belief-action deviations as structured evidence, which aids in understanding agent decisions.
  • Active-belief conditions improve outcomes, but the direct action-belief consistency is low, suggesting complex underlying mechanisms.

Possible limitations

  • The mechanism behind improved outcomes under active-belief conditions remains unresolved.
  • Not stated in the abstract

Abstract

arXiv:2607.10814v1 Announce Type: cross Abstract: Evaluating LLM agents in hidden-information multi-agent settings is hard: final outcomes are high-variance and rarely reveal why an agent decided as it did. We study this in a 9-player Werewolf environment where agents act under strict, code-level information isolation, and we build an auditable framework that maintains an external belief state over hidden roles, logs belief updates and belief-action deviations as structured evidence, and supports a defensive offline improvement loop that reviews bad cases before any strategy change. Across 1,080 frozen games spanning belief-disabled, active-belief, kernel-ablation, camp-restricted, consumption-policy, and high-load arms, and including a seed-paired A0/A1 comparison, the active-belief condition is associated with substantially better good-side outcomes: in the 200-seed A0/A1 comparison the good-side win rate rises from 0.205 to 0.390 (paired McNemar $\chi^2 = 16.4$, $p < 0.001$), with fewer irreversible witch-poison errors. We do not, however, attribute this shift to belief content. Direct action-belief consistency is low ($\approx 0.21$), and giving belief only to the werewolves helps the good side more than giving it only to the good side, which argues against a simple holder-benefit account; we therefore report the effect as an association and treat its mechanism as unresolved. The contribution is the audit framework itself: it makes the effect measurable, exposes low direct action-belief consistency, rejects an unreliable forced-consumption intervention with evidence, and separates strategy effects from load confounds. We accordingly position external belief in high-noise hidden-information games primarily as an auditable cognitive baseline that also carries decision-relevant signal, turning opaque agent behavior into replayable evidence for safer, controlled iteration.