Back to today's list

Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure

Jianshuo Dong, Yiming Liu, Maosen Zhang, Nan Deng, Xu Peng, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu

Published Aug 5, 2026Featured #2In the daily list Aug 6, 2026
Daily score73.1
Editorial review7.5
Relevance0.461
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding and mitigating vulnerabilities in LLMs is crucial for ensuring their safe deployment in real-world applications, making this research relevant for both developers and researchers in AI security.

This research reveals how agentic LLMs can be probed for indirect prompt injection vulnerabilities and introduces a novel defense mechanism.

Summary

The paper investigates the vulnerability of agentic LLMs to indirect prompt injection (IPI) attacks, analyzing how these models encode signals of IPI exposure and proposing a defense mechanism called AGRI that enhances safety while maintaining task utility.

Key contributions

  • Development of probing techniques that achieve high AUROC scores for predicting IPI exposure in LLMs.
  • Introduction of the AGRI defense mechanism that significantly reduces attack success rates while preserving model performance.
  • Creation of an analysis framework for correlating natural-language explanations with probe-captured signals.

Notable insights

  • The introduction of a recognition-action gap highlights a critical area where LLMs fail to act on encoded signals, which could inform future safety mechanisms.
  • The use of linear probes to predict IPI exposure across multiple models demonstrates a practical approach to evaluating model vulnerabilities.

Possible limitations

  • Not stated in the abstract.

Abstract

arXiv:2608.02657v1 Announce Type: cross Abstract: Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While many efforts have sought to address the threats, little is known about the internals of agentic LLMs when they are exposed to IPI attacks, a condition which we call IPI exposure. In this paper, we study this problem in depth from three aspects. (1) Probing: Across six models, including the giant 753B-parameter GLM-5.2, simple linear probes trained on pre-generation hidden states can predict LLMs' IPI exposure. These probes achieve 90%+ AUROC on unseen attacks, agent instructions, and task suites; they exhibit high robustness under adaptive attacks and in cross-lingual settings. (2) Defense: Our CoT measurement reveals a recognition--action gap: though models encode such signals, they often fail to translate them into safe actions. We then introduce AGRI, a probe-gated reasoning-based defense that prepends anti-injection reasoning on demand. On difficult AgentDojo settings, AGRI substantially reduces attack success rate, e.g., from 34.6% to 0% on Qwen3.5-27B, while largely maintaining clean-task utility. (3) Explanation: We introduce an analysis framework that identifies natural-language explanations most strongly correlated with probe-captured signals. The resulting profiles differ across models: latent signals can align with either direct IPI-exposure claims or indirect operational cues. Code is available: https://github.com/jianshuod/IPI-exposure-signal.