Back to today's list

Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways

Shuyi Miao, Wangjie Qiu, Pengyang Shao, Canran Xiao, Fei Shen, Zhiming Zheng, Tat-Seng Chua

Published Oct 1, 2026
Editorial review6.8
Relevance0.485
Freshness0.000

Why It Matters

What makes this one worth your time

Understanding and improving cross-lingual safety mechanisms in LLMs is crucial for developing AI systems that are safe and reliable across diverse languages, especially those with fewer resources.

The paper uncovers cross-lingual pathways in LLMs that enhance safety in low-resource languages.

Summary

The paper investigates the internal mechanisms of large language models (LLMs) to understand how safety capabilities are shared across languages. It identifies cross-lingual shared safety pathways that transfer safety capabilities from high-resource languages to non-high-resource languages and proposes a targeted alignment method to improve safety in these languages.

Key contributions

  • Identification of monolingual and cross-lingual safety pathways in LLMs.
  • Proposal of a pathways-targeted alignment method to improve safety in non-high-resource languages.

Notable insights

  • Cross-lingual shared safety pathways act as bridges for transferring safety capabilities between languages.
  • Targeting specific pathway parameters can enhance safety in non-high-resource languages without compromising overall model performance.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2608.09095v2 Announce Type: replace Abstract: Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy artificial intelligence. Currently, mechanistic interpretability studies on multilingual safety are largely confined to local components, such as isolated neurons. However, this static and fragmented perspective overlooks the synergy among components and fails to elucidate how safety signals dynamically propagate within the model to drive safety decisions ultimately. In this work, we move beyond isolated neurons to identify and target the cross-layer functional pathways formed during safety signal propagation, thereby uncovering the mechanisms driving the cross-lingual safety gap. Specifically, we first identify monolingual safety pathways and validate their impact on refusing harmful requests. Subsequent cross-lingual analyses reveal a sparse subset of cross-lingual shared safety pathways, confirming that this intersection acts as the internal bridge transferring safety capabilities from high-resource (HR) languages to non-high-resource (NHR) languages. Building on these mechanistic findings, we propose a pathways-targeted alignment method based on the cross-lingual shared safety pathways. Experimental results show that updating only a small fraction of pathway parameters significantly improves safety in NHR languages while largely preserving the model's general capabilities.