Back to today's list

Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents

Minghui Pan, Jiayuxuan Yang, Yuanyuan Yuan, Yu Jiang, Zhenpeng Chen

Published Aug 4, 2026Featured #4In the daily list Aug 5, 2026
Daily score69.0
Editorial review7.2
Relevance0.513
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding and mitigating safety risks in AI agents is crucial for deploying them in real-world applications where safety is a priority.

SafeKeep enhances AI agent safety by decoupling safety judgment from tool execution.

Summary

The paper identifies schema-formatted tool specifications as a key factor in the safety degradation of AI agents and introduces SafeKeep, a safeguard that improves safety by decoupling safety judgment from tool execution. SafeKeep significantly increases refusal rates for harmful requests and reduces attack success rates across multiple benchmarks and models.

Key contributions

  • Identification of schema-formatted tool specifications as a source of safety degradation.
  • Development of SafeKeep, an inference-time safeguard that improves safety by decoupling safety judgment from tool execution.
  • Demonstration of SafeKeep's effectiveness across multiple benchmarks and models.

Notable insights

  • Schema-formatted tool specifications can weaken internal refusal signals in AI agents.
  • Decoupling safety judgment from tool execution can significantly improve safety outcomes.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2607.29254v1 Announce Type: new Abstract: AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into consequential real-world actions. Yet LLMs often become substantially less safe when deployed as agents, and the source of this degradation remains poorly understood. In this paper, we identify schema-formatted tool specifications as a primary source of agent safety degradation and show, through white-box representation analysis, that they weaken the model's internal refusal signals and contribute to unsafe tool execution. Building on this finding, we propose SafeKeep, an inference-time safeguard that decouples safety judgment from tool execution: it assesses requests using flattened textual tool specifications while retaining the original schema-formatted specifications for execution. Across two representative benchmarks and four LLMs, including both white-box and black-box models, SafeKeep increases the average refusal rate for harmful requests from 23.8% to 70.6% and reduces the average attack success rate under observation-level prompt injection from 25.6% to 2.5%. It also outperforms existing safeguards and preserves task-handling capability. We release the code and data at https://github.com/snowcatsmoking/SafeKeep .