Back to today's list

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

Youting Wang, Xiao Han, Dingyan Shang, Yuan Tang, Bowen Liu

Published Aug 4, 2026Featured #6In the daily list Aug 5, 2026
Daily score62.2
Editorial review6.8
Relevance0.500
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding the validity of safety benchmarks is crucial for developing reliable AI systems, as it impacts how safety is measured and interpreted in AI models.

The paper critically examines the validity of agent-safety benchmarks, highlighting inconsistencies in their measurements.

Summary

The paper evaluates four agent-safety benchmarks (R-Judge, InjecAgent, AgentHarm, AgentDojo) by testing them on up to 22 models to assess their validity in measuring agent safety, revealing inconsistencies and correlations between capability and safety metrics.

Key contributions

  • Validation of four agent-safety benchmarks across multiple models.
  • Identification of inconsistencies in benchmark rankings and their implications for safety measurement.

Notable insights

  • The paper identifies that different benchmarks rank models inconsistently, suggesting that benchmark choice significantly affects perceived safety.
  • It reveals a negative correlation between capability and misalignment safety, indicating potential trade-offs in model design.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2607.28685v1 Announce Type: new Abstract: Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentDojo) as measurements to be validated, running each under its official implementation and author-provided scorer on up to 22 models, with MMLU and GPQA measured by us under one protocol as a capability composite. The metric is the first problem. On any binary trace-judgment benchmark scored by $F_1$, an ``always positive'' policy attains $F_1 = 2\pi/(1+\pi)$; on R-Judge that is $0.690$, above five of the 21 models that actually discriminate. The three broad-coverage benchmarks then rank the same 18 models differently, and the trade-off behind that disagreement is a small-panel artifact: R-Judge specificity against AgentHarm safety correlates $-0.64$ at $n{=}7$ and $+0.02$ at $n{=}18$, and a quarter of random size-7 subsets reach $|\rho| \geq 0.5$ around that near-zero value. Held-out validity turns on which outcome you pick. Capability predicts task success ($\rho{=}{+}0.60$) but correlates negatively with misalignment safety ($\rho{=}{-}0.44$, $n{=}21$). On their paired $n{=}20$ panel, the corresponding contrast is $\Delta{=}{-}1.00$ (95% CI $[-1.48, -0.49]$, $p<0.001$), and it survives leave-one-organization-out and organization-clustered bootstrap analyses. On an expanded 41-model panel, the misalignment correlation weakens to $-0.16$ (95% CI $[-0.54, +0.22]$) and jailbreak strengthens to $+0.34$, though neither change is significant. \mbox{AgentHarm} shows the strongest held-out association, $\rho{=}{+}0.72$ with three-template jailbreak safety after controlling capability. But both instruments score harmful compliance, so this is evidence of convergent validity rather than general safety. Naming the benchmark, metric, target behavior, and model panel is the minimum a safety claim needs.