Reason-Mediated Behavioral Models for Auditing LLM Social Simulators
Atharva Pandey, Gautam Jajoo
Why It Matters
What makes this one worth your time
Understanding the alignment between LLM-generated rationales and human reasoning is crucial for improving the reliability and interpretability of AI systems used in social simulations.
The paper proposes a framework to audit LLM social simulators by comparing human and simulated rationale-derived reasons.
Summary
The paper introduces a framework for evaluating large language models (LLMs) used as social simulators by comparing human rationale-derived reasons with LLM-simulated reasons in predicting behavior. It highlights the brittleness of LLM-generated reasons, which often mimic the concept board rather than accurately reflecting human decision paths.
Key contributions
- Proposes an evaluation framework for auditing LLM social simulators.
- Introduces the concept of signed reason states to compare human and LLM rationales.
Notable insights
- Mapping human rationales into signed reason states provides a practical audit tool for evaluating LLMs.
- LLM-simulated reasons often echo the concept board rather than accurately capturing human reasoning paths.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2607.24649v1 Announce Type: new Abstract: Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether simulated outcomes resemble human outcomes. We argue that this is necessary but too weak: a simulator can match the final answer while using the wrong rationale-derived reason pattern. We study this problem through a 94-person sunscreen concept test in which each respondent evaluated three product concepts and wrote open-ended rationales. We map those rationales into signed reason states $Z$, where positive signs support adoption and negative signs block it. This gives a practical audit: holding respondent descriptors $D$, category context $K$, and concept treatment $X$ fixed, do human rationale-derived reasons help predict behavior $Y$, and can an LLM simulate the same reason state without seeing the human rationale or outcome? Human rationale-derived reasons substantially improve held-out prediction of purchase intent. LLM-simulated reasons are more brittle: they often sound plausible, but frequently echo the concept board rather than recover the respondent's acceptance or rejection path. The paper contributes an evaluation framework for social simulators. Reason states do not identify natural causal effects by themselves, but they provide an interpretable test of whether a simulator's stated reasons align with human evidence.