Back to today's list

GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory

Pepijn Cobben, Xuanqiang Angelo Huang, Thao Amelia Pham, Isabel Dahlgren, Terry Jingchen Zhang, Zhijing Jin

Published May 25, 2026Featured #4In the daily list May 26, 2026
Daily score70.3
Editorial review7.5
Relevance0.452
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding multi-agent risks is crucial for developing safer AI systems, especially in high-stakes environments where coordination and conflict can have serious consequences.

GT-HarmBench provides a critical benchmark for assessing AI safety in multi-agent contexts.

Summary

The paper introduces GT-HarmBench, a benchmark designed to evaluate AI safety risks in multi-agent environments using game-theoretic scenarios, revealing significant failures in socially beneficial actions among agents.

Key contributions

  • Development of GT-HarmBench as a standardized testbed for multi-agent AI safety evaluation.
  • Empirical analysis showing a 38% failure rate in socially beneficial actions across frontier models.
  • Demonstration of the effectiveness of game-theoretic interventions in improving outcomes.

Notable insights

  • The benchmark includes a diverse set of 1,535 scenarios that reflect realistic AI risk contexts.
  • Game-theoretic interventions demonstrated a measurable improvement in socially beneficial outcomes.

Possible limitations

  • Not stated in the abstract.

Abstract

arXiv:2602.12316v2 Announce Type: replace Abstract: Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely evaluate single agents, leaving multi-agent risks such as coordination failure and conflict poorly understood. We introduce GT-HarmBench, a benchmark of 1,535 high-stakes scenarios spanning game-theoretic structures such as the Prisoner's Dilemma, Stag Hunt and Chicken. Scenarios are drawn from realistic AI risk contexts in the MIT AI Risk Repository. Across 15 frontier models, agents fail to choose socially beneficial actions in 38% of high-stakes cases, such as military escalation, election manipulation, and medical malpractice. We measure sensitivity to game-theoretic prompt framing and ordering, and analyze reasoning patterns driving failures. We further show that game-theoretic interventions improve socially beneficial outcomes by up to 18%. Our results highlight substantial reliability gaps and provide a broad standardized testbed for studying alignment in multi-agent environments. The benchmark and code are available at https://github.com/causalNLP/gt-harmbench.