EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures
Bu\u{g}ra Alperen Ulu{\i}rmak, Rifat Kurban
Why It Matters
What makes this one worth your time
Understanding and improving the evaluation and safety of LLMs is crucial for developing reliable AI systems, making this framework potentially valuable for researchers and practitioners focused on AI safety and evaluation.
EvalSafetyGap offers a new framework for understanding evaluation-safety failures in LLMs.
Summary
The paper introduces EvalSafetyGap, a hybrid survey and conceptual framework aimed at addressing evaluation-safety failures in large language models (LLMs). It combines a systematic search with narrative synthesis and a structured audit of ten models, covering various evidence streams related to LLM evaluation and safety. The study highlights the challenges in measuring capability, behavioral safety, and governance separately, and proposes new constructs like Instability Decomposition and Alignment Trilemma to facilitate testable comparisons.
Key contributions
- Introduction of the EvalSafetyGap framework for comparing evaluation-side and alignment-side proxy failures.
- Development of Instability Decomposition and Alignment Trilemma constructs.
- A structured audit of ten models to assess the relationship between capability and adversarial robustness.
Notable insights
- The use of Instability Decomposition and Alignment Trilemma as tools for generating testable comparisons is a novel approach.
- The audit reveals that governance and disclosure play a more significant role in safety gaps than behavioral robustness.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2606.30219v4 Announce Type: replace Abstract: This paper presents a systematic survey and conceptual synthesis of the shared measurement problem underlying large language model (LLM) evaluation and AI safety: benchmark scores, reward signals, and safety metrics can improve while the capabilities and alignment properties they are meant to represent remain uncertain. Synthesizing 373 primary studies published between 2018 and 2026, the survey organizes evidence on benchmark validity, contamination, dynamic evaluation, LLM-as-a-judge protocols, adversarial safety testing, reward and proxy optimization, mechanistic interpretability, and AI governance into an eight-stream evidence taxonomy. Building on this synthesis, we introduce EvalSafetyGap, a conceptual framework that unifies benchmark-validity and alignment-failure research as a shared proxy-target divergence problem under optimization pressure, formalized through a Goodhart-inspired Instability Decomposition and an Alignment Trilemma. An exploratory ten-model public-evidence audit illustrates the framework by showing why capability, behavioral robustness, and governance disclosure should be reported as separate evidence layers rather than collapsed into a single safety score. The survey closes with a research agenda for dynamic and contamination-resistant benchmarks, pre-specified multi-attempt threat models, version-locked evaluation, transparent source reporting, and validated mechanistic safety indicators, offering researchers, model developers, and AI auditors a shared vocabulary for measurement-aware LLM safety evaluation.