Back to today's list

AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models

Rohan Naphade, Minzhou Pan, Bo Li

Published Jul 28, 2026
Editorial review6.8
Relevance0.469
Freshness0.000

Why It Matters

What makes this one worth your time

As AI models and regulations rapidly evolve, a dynamic benchmark like AIR-BENCH Live helps ensure ongoing safety and compliance, which is crucial for developers and policymakers.

AIR-BENCH Live is a self-updating safety benchmark for foundation models that adapts to new regulations.

Summary

The paper introduces AIR-BENCH Live, a dynamic safety benchmark for foundation models that evolves by integrating new government regulations and generating updated prompts using a multi-agent, persona-driven algorithm.

Key contributions

  • Development of a self-evolving safety benchmark for foundation models.
  • Implementation of an automated pipeline to integrate new regulations into the benchmark.
  • Introduction of a multi-agent, persona-driven prompt generation algorithm.

Notable insights

  • The use of a multi-agent, persona-driven prompt generation algorithm to create realistic and multilingual prompts.
  • The automated update pipeline that classifies new policies against an existing risk taxonomy and proposes new categories.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2607.22671v1 Announce Type: new Abstract: Foundation-model safety benchmarks capture the AI risks of their time of publication: as models improve and governments pass new AI-safety legislation, their risk taxonomies become incomprehensive and their attack prompts become ineffective. We present AIR-BENCH Live, a self-evolving successor to AIR-BENCH 2024. An automated update pipeline monitors government regulation and classifies new policies against the current four-tier risk taxonomy, either matching them to existing categories or proposing new granular categories. Then, a multi-agent, persona-driven prompt generation algorithm generates realistic, multilingual prompts with minimal human review, leaving room for improvement with modern jail breaking techniques. This algorithm is used to overhaul legacy prompts and generate prompts for new categories. In our current version, the pipeline has expanded the benchmark from 314 to 335 granular risks, with the 21 new categories drawing from 31 truly novel policy clauses across seven jurisdictions. Evaluating 14 recent models, we find a wide safety spread (from 0.17 to 0.89 among the models judged on their own behavior), that the modernized prompts are on average 0.06 points harder than the 2024 set, with the largest drops concentrated among the most compliant models, and that most models are modestly less safe on non-English prompts. By continuously absorbing new regulation and regenerating prompts, AIR-BENCH Live is designed to evolve alongside a fast-moving field.