Back to today's list

$S^3$: Improving Agent Safety through Multi-Stage Defense

Zibo Xiao, Haoyu Wang, Jun Sun

Published Aug 5, 2026
Editorial review6.8
Relevance0.484
Freshness0.000

Why It Matters

What makes this one worth your time

Ensuring safety in complex agent workflows is crucial for reliable AI deployment, and this framework offers a structured approach to address risks across different stages.

A framework for improving agent safety through stage-specific safety skills.

Summary

The paper proposes a framework called $S^3$ that enhances agent safety by integrating stage-specific safety skills into multi-stage agentic workflows. It introduces a guard agent to orchestrate these skills for risk detection and mitigation, and evaluates the framework using a new Multi-Stage Risk Benchmark (MSRB).

Key contributions

  • Introduction of Stage-Specific Safety Skills as reusable components.
  • Development of an automated transformation pipeline for safety designs.
  • Proposal of the Multi-Stage Risk Benchmark (MSRB) for evaluating risks.

Notable insights

  • The use of a guard agent to orchestrate safety skills across workflow stages is a novel approach to comprehensive risk management.
  • The creation of a community-driven safety skill library suggests a collaborative effort to improve agent safety.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2608.02683v1 Announce Type: cross Abstract: Large Language Model (LLM) agents rely on multi-stage agentic workflows, with stages such as memory, planning, and tool execution, to accomplish complex tasks. However, risks may emerge at different stages, propagate across steps, and become difficult to detect and mitigate. Existing safety methods protect only isolated stages and are difficult to integrate, leaving agents without comprehensive protection throughout the workflow. To address these limitations, we introduce Stage-Specific Safety Skills, a unified abstraction that represents heterogeneous safety designs as reusable and composable components with explicit stage semantics. We further develop an automated transformation pipeline that converts existing safety designs into reusable safety skills and establish a community-driven safety skill library. Building on this abstraction, we propose $S^3$, a multi-stage defense framework in which a guard agent orchestrates stage-specific safety skills for risk detection and mitigation throughout the agentic workflow. We also construct the Multi-Stage Risk Benchmark (MSRB) to evaluate representative risks across workflow stages. Experimental results show that $S^3$ consistently outperforms representative state-of-the-art baselines in both safety effectiveness and utility preservation. These results demonstrate the potential of stage-specific safety skills as a scalable and composable foundation for building resilient and trustworthy agent systems.