Back to today's list

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

Enyi Jiang, Anders Gj{\o}lbye, Yibo Jacky Zhang, Sanmi Koyejo

Published Aug 5, 2026Featured #3In the daily list Jun 10, 2026
Daily score72.9
Editorial review7.5
Relevance0.470
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding representation-level vulnerabilities is crucial for developing more robust AI systems, as current evaluations may provide a false sense of security.

This research reveals critical vulnerabilities in LLMs that behavioral safety metrics overlook.

Summary

The paper identifies a gap between behavioral safety evaluations and representation-level robustness in large language models, proposing a new evaluation framework and the Latent Vulnerability Score (LVS) to assess model vulnerabilities under intervention.

Key contributions

  • Formalization of the audit gap between behavioral safety and representation-level robustness.
  • Development of an intervention-based evaluation framework for testing model robustness.
  • Introduction of the Latent Vulnerability Score (LVS) for measuring latent vulnerabilities.

Notable insights

  • The introduction of the Latent Vulnerability Score (LVS) offers a novel metric for assessing model robustness beyond observable behavior.
  • Dissociated models can maintain safe outward behavior while being internally vulnerable, highlighting the need for deeper evaluations.

Possible limitations

  • Not stated in the abstract.

Abstract

arXiv:2606.08044v2 Announce Type: replace-cross Abstract: Safety evaluation of large language models (LLMs) is largely behavioral: a model is certified safe when it refuses harmful requests and answers benign ones. But refusing on the prompts an auditor happens to try does not show that the model is far from harmful behavior. Behavioral tests observe outputs; they do not measure how easily an intervention on the model turns a refusal into compliance. We call the gap between what static audits certify and what an intervention can reach the audit gap, and we show it is realizable: one can build a model that matches its safety-aligned base on every static audit yet gives way to a small, known perturbation of its internal state. We construct such dissociated models from three safety-aligned bases (Gemma 2 2B, Llama 3.2 3B, Qwen 2.5 3B) and audit the base, dissociated, and openly harmful models with the same soft interventions in parameter and latent space; the latent attacks are summarized by the Latent Vulnerability Score (LVS), the safety degradation produced per unit of bounded latent perturbation. Every static audit we run gives the dissociated model the same verdict as its base, since its refusals match the base, jailbreaks show no consistent signature, and a strong fixed probe on clean activations cannot tell it from the base. The same interventions an auditor could run reverse the verdict. At the targeted mid layer the dissociated models score 2.5 to 3.1 times higher LVS than their bases. A bounded latent attack elicits harmful compliance on 54 to 86% of prompts, against 3 to 48% for the bases, while matched random perturbations stay at or below 12%. Harmful fine-tuning reaches high compliance within five gradient steps, where the bases need 10 to 25. Behavioral testing, even with static latent probing, cannot certify representation-level robustness: a safety audit must intervene on the model, not only observe it.