Back to today's list

When Words Are Safe But Actions Kill: Probing Physical Jailbreak Beyond Textual Jailbreak in Hidden-State Risk Space

Weimeng Wang, Ziqiang Wang, Zihang Zhan, Chuanpu Fu, Qi Li, Ke Xu

Published Aug 24, 2026Featured #7In the daily list Jul 19, 2026
Daily score69.8
Editorial review7.5
Relevance0.458
Freshness0.722

Why It Matters

What makes this one worth your time

Understanding the difference between text-based safety and real-world risks is crucial for the safe deployment of AI systems in physical environments.

This research reveals how language model instructions can lead to physical dangers, proposing a novel detection method.

Summary

The paper investigates the distinction between content danger and physical danger in large language models, proposing a method called PRISM to effectively identify physical risks in grounded actions, supported by experimental results on new benchmarks.

Key contributions

  • Development of PRISM, a logistic probe for detecting physical danger in LLMs.
  • Introduction of the PhysicalSafetyBench-1K benchmark for evaluating physical risk detection.
  • Demonstration of the separability of content danger and physical danger in LLM representations.

Notable insights

  • The study introduces a hidden-state direction analysis to differentiate between content danger and physical danger in LLMs.
  • The creation of the PhysicalSafetyBench-1K benchmark allows for a more nuanced evaluation of physical risk detection beyond traditional text moderation.

Possible limitations

  • Not stated in the abstract.

Abstract

arXiv:2607.15218v2 Announce Type: replace Abstract: Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world. We study whether this physically grounded jailbreak is the same safety problem as ordinary textual jailbreak. Through hidden-state direction analysis and random-split null tests, we show that textual jailbreak (TJ) and physical jailbreak (PJ) form separable signals in LLM representations across Qwen2.5-3B/7B/14B/32B, Phi-3.5 and SmolLM2. Building on this separability, we propose PRISM, a single-layer L2-regularized logistic probe over full hidden states. PRISM achieves 86.2--87.7\% accuracy on SafeAgentBench with 11.7--13.7\% false-positive rates (FPRs), while same-scale LLM judges over-block safe tasks at 24.7--39.0\% FPR. To test whether the result survives lexical-shortcut controls, we introduce an interaction-balanced revision of PhysicalJailbreakBench-2K (PJB-2K): a fixed 2{,}000-row comparison set sampled by label and physical mechanism from a larger object--site construction. On the underlying 10{,}000-row pool, word-TFIDF and the embedding layer remain at chance (AUC 0.497 and 0.500). At layer 25, selected by an i.i.d. sweep, cell-grouped cross-validation gives PRISM 0.718 AUC, compared with 0.398 for a physics-free label control under the same protocol. On the identical 2{,}000 comparison rows, these PRISM predictions obtain 0.671 balanced accuracy, while Qwen2.5 judges from 3B to 72B obtain 0.538--0.577 and exhibit high FPR. These results support hidden-state probing as a representation-level method for physical safety beyond text moderation, without relying on the near-perfect scores of shortcut-prone paired templates.