Back to today's list

Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning

Ian B. de Haan, Peter van der Putten, Max van Duijn

Published Aug 6, 2026
Editorial review6.8
Relevance0.457
Freshness0.000

Why It Matters

What makes this one worth your time

Understanding the robustness of reasoning models can inform future developments in AI systems that require nuanced understanding of human-like reasoning.

This study highlights the robustness of reasoning models in Theory of Mind tasks over traditional capabilities.

Summary

The paper investigates the performance of reasoning-oriented large language models on Theory of Mind tasks, revealing that these models show increased robustness to variations in prompts and tasks, suggesting a robustness-based explanation rather than a new ToM-specific ability.

Key contributions

  • Analysis of reasoning models' performance on Theory of Mind tasks.
  • Demonstration of robustness in reasoning models under prompt and task variations.
  • Introduction of novel adaptations of psychological experiments for machine evaluation.

Notable insights

  • The adaptation of machine psychological experiments for evaluating ToM in LLMs is a novel methodological approach.
  • The findings suggest that improvements in reasoning models may not indicate enhanced ToM abilities but rather a general robustness to variations.

Possible limitations

  • Not stated in the abstract.

Abstract

arXiv:2608.04646v1 Announce Type: new Abstract: Large language models (LLMs) have recently shown strong performance on Theory of Mind (ToM) tests, prompting debate about the nature and validity of the underlying capabilities. At the same time, reasoning-oriented LLMs trained via reinforcement learning with verifiable rewards have demonstrated notable improvements across a range of benchmarks. In this work, we examine the behavior of such reasoning models in ToM tasks using novel adaptations of machine psychological experiments together with results from established benchmarks. We observe that reasoning models consistently exhibit increased robustness to prompt variations and task perturbations. Our analysis suggests these gains come at least partly from models being more robust at reaching the correct answer under prompt and task variation. We read this as evidence for a robustness-based account rather than for a new ToM-specific ability.