Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks
Shangze Li, Chuancheng Shi, Simiao Xie, Lingzhi He, Cheng Ji, Zifeng Cheng, Fei Shen, Chao Wu, Tat-Seng Chua
Why It Matters
What makes this one worth your time
As large foundation models become more prevalent, ensuring their safety against sophisticated attacks is crucial for maintaining trust and reliability in AI systems.
DRAA offers a novel approach to fortify model safety against targeted white-box attacks.
Summary
The paper proposes a framework called dynamic routing adaptive alignment (DRAA) to enhance the robustness of large foundation models against white-box attacks by introducing dynamic compensatory routes that adapt when safety routes are compromised.
Key contributions
- Introduction of the DRAA framework for dynamic routing in model safety.
- Methodology for identifying and localizing safety routes using internal activations.
- Construction of failure-aware preference pairs to enhance robustness against attacks.
Notable insights
- The approach of dynamically masking safety routes to induce failure cases is a clever way to identify vulnerabilities.
- The method of contrasting internal activations between safe and unsafe samples provides a novel diagnostic tool for safety route localization.
Possible limitations
- Not stated in the abstract.
Abstract
arXiv:2608.02674v1 Announce Type: cross Abstract: With the widespread deployment of large foundation models (LFMs) in open environments, safety threats are shifting from black-box jailbreaks toward white-box attacks that directly identify and disrupt internal safety neurons or routes. However, existing safety defenses often rely on static safety units or fixed refusal pathways, leaving models highly vulnerable to targeted route-level white-box attacks. For that, we propose dynamic routing adaptive alignment (DRAA), a framework that introduces dynamic compensatory routes to preserve robust refusal behavior when the safety route is compromised. Specifically, we first identify and localize the model's safety route by contrasting internal activations between safe and unsafe calibration samples. DRAA then masks this safety route to induce causal failure cases and selectively mines the resulting defense failures, thereby constructing failure-aware preference pairs. Extensive experiments demonstrate that DRAA effectively restructures the underlying pathway dependence of model safety, substantially improving robustness against route-level white-box attacks, while preserving general utility.