Back to today's list

When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas

Steffen Backmann, David Guzman Piedrahita, Terry Jingchen Zhang, Emanuel Tewolde, Rada Mihalcea, Bernhard Sch\"olkopf, Zhijing Jin

Published Jul 27, 2026Featured #9In the daily list Jul 28, 2026
Daily score56.7
Editorial review6.8
Relevance0.491
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding how LLMs navigate ethical dilemmas is crucial for their safe deployment in real-world scenarios where moral and profit incentives may conflict.

The paper evaluates LLMs' moral behavior in social dilemmas where ethics and profit incentives conflict.

Summary

The paper introduces a framework, msimfull (msim), to evaluate the behavior of large language models (LLMs) in morally charged social dilemmas like the prisoner's dilemma and public goods game. It examines how these models balance moral imperatives against profit incentives, analyzing factors such as moral framing and opponent behavior. The study finds variability in moral behavior across models and identifies game structure and moral framing as key drivers of behavior.

Key contributions

  • Introduction of msimfull (msim) framework for evaluating LLMs in morally charged social dilemmas.
  • Analysis of causal drivers of moral behavior using average treatment effects.
  • Characterization of motive profiles through reasoning-trace analysis.

Notable insights

  • The study uses average treatment effects (ATEs) to estimate the causal impact of different factors on LLM behavior.
  • Reasoning-trace analysis reveals distinct motive profiles across models, highlighting variability in moral versus payoff-maximizing behavior.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2505.19212v2 Announce Type: replace Abstract: Recent advances in LLMs have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical alignment a critical concern. While prior work has examined LLMs' moral judgment and strategic behavior separately, there is limited understanding of how they act when moral imperatives directly conflict with profit incentives. We introduce \msimfull (\msim) to evaluate how LLMs behave in the prisoner's dilemma and public goods game embedded in morally charged contexts, varying moral framing, opponent behavior, and survival pressure across nine models. Beyond measuring behavior, we estimate the causal effect of each factor via average treatment effects (ATEs) and analyze agents' own reasoning traces to characterize the motives behind their choices. We find that no model remains consistently moral, with cooperation rates ranging from 7.9\% to 76.3\%. Game structure and moral framing are the strongest causal drivers of moral behavior, while reasoning-trace analysis reveals distinct motive profiles across models, ranging from predominantly payoff-maximizing to moral- and reputation-oriented. Together, these results expose the situational brittleness of current LLMs' moral behavior and the risk of deploying them where profit incentives conflict with ethical guidelines.