Back to today's list

Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind

Hanqi Xiao, Vaidehi Patil, Zaid Khan, Hyunji Lee, Elias Stengel-Eskin, Mohit Bansal

Published Jul 24, 2026Featured #6In the daily list Jul 25, 2026
Daily score60.8
Editorial review6.8
Relevance0.474
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding and improving AI's ability to model and manipulate beliefs is crucial for developing robust systems capable of safe interactions in adversarial settings.

The paper proposes a new ToM-based challenge for AI models to act as Double Agents in belief manipulation scenarios.

Summary

The paper introduces a novel challenge called ToM for Steering Beliefs (ToM-SB), where a defender acts as a Double Agent to manipulate an attacker's beliefs using theory of mind (ToM) in a shared universe. The study evaluates the performance of AI models trained with reinforcement learning on this task, showing that combining ToM and fooling rewards enhances performance against various attackers, outperforming existing models like Gemini3-Pro and GPT-5.4.

Key contributions

  • Introduction of the ToM for Steering Beliefs (ToM-SB) challenge.
  • Demonstration of improved performance through reinforcement learning with combined ToM and fooling rewards.
  • Evaluation of AI Double Agents against multiple attacker models, showing generalization to out-of-distribution settings.

Notable insights

  • There is a bidirectional relationship between theory of mind and attacker-fooling capabilities.
  • Combining ToM and fooling rewards leads to superior performance in belief manipulation tasks.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2604.11666v2 Announce Type: replace-cross Abstract: As large language models (LLMs) become the engine behind conversational systems, their ability to reason about the intentions and states of their dialogue partners (i.e., form and use a theory-of-mind, or ToM) becomes increasingly critical for safe interaction with potentially adversarial partners. We propose a novel privacy-themed ToM challenge, ToM for Steering Beliefs (ToM-SB), in which a defender must act as a Double Agent to steer the beliefs of an attacker with partial prior knowledge within a shared universe. To succeed on ToM-SB, the defender must engage with and form a ToM of the attacker, with a goal of fooling the attacker into believing they have succeeded in extracting sensitive information. We find that strong frontier models like Gemini3-Pro and GPT-5.4 struggle on ToM-SB, often failing to fool attackers in hard scenarios with partial attacker prior knowledge, even when prompted to reason about the attacker's beliefs (ToM prompting). To close this gap, we train models on ToM-SB to act as AI Double Agents using reinforcement learning, testing both fooling and ToM rewards. Notably, we find a bidirectionally emergent relationship between ToM and attacker-fooling: rewarding fooling success alone improves ToM, and rewarding ToM alone improves fooling. Across four attackers with different strengths, six defender methods, and both in-distribution and out-of-distribution (OOD) evaluation, we find that gains in ToM and attacker-fooling are well-correlated, highlighting belief modeling as a key driver of success on ToM-SB. AI Double Agents that combine both ToM and fooling rewards yield the strongest fooling and ToM performance, outperforming Gemini3-Pro and GPT-5.4 with ToM prompting on hard scenarios. We also show that ToM-SB and AI Double Agents can be extended to stronger attackers, demonstrating generalization to OOD settings and the upgradability of our task.