ForesightSafety-SAGE:A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents
Lu Jia, Haibo Tong, Feifei Zhao, Jindong Li, Dongqi Liang, Ping Wu, Qian Zhang, Yi Zeng
Why It Matters
What makes this one worth your time
Understanding and improving the safety of LLM agents is crucial as they become more autonomous and integrated into real-world applications.
ForesightSafety-SAGE automates scenario generation and safety evaluation for LLM agents, highlighting significant safety risks.
Summary
The paper introduces ForesightSafety-SAGE, an automated framework for generating scenarios and evaluating the safety of large language model agents across five risk dimensions, resulting in 1,072 evaluation scenarios. It evaluates 12 LLM agents, revealing significant behavioral safety risks during task execution.
Key contributions
- Development of an automated scenario generation framework for LLM safety evaluation.
- Creation of 1,072 measurable evaluation scenarios based on five risk dimensions.
- Evaluation of 12 LLM agents, highlighting substantial safety risks.
Notable insights
- Automated scenario generation can provide a more comprehensive safety evaluation than static prompts or manual scenarios.
- Evaluating agents under different authority contexts can reveal varying levels of safety risks.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2606.08531v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, access external environments, and execute tasks. As their capabilities and autonomy expand, the safety risks they face also become more diverse. Existing evaluations often rely on manually written scenarios, static prompts, or final-output judgments, making it difficult to capture the diverse risks that agents may face during task execution. We introduce ForesightSafety-SAGE, a fully automated scenario generation and safety evaluation framework for LLM agents. Based on five risk dimensions,we instantiae abstract and diverse safety risks in real-world task execution into 1,072 measurable evaluation scenarios. Using the automated evaluation pipeline, 12 LLM agents are evaluated under two authority contexts. The results show that current agents still face substantial behavioral safety risks during task execution, with an average ASR of 47.1% and several models exceeding 70%. These findings demonstrate the importance of executable, process-level evaluation for understanding and improving LLM agent safety.