Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment
Alina Klerings, Jannik Brinkmann, Heiner Stuckenschmidt, Simone Paolo Ponzetto
Why It Matters
What makes this one worth your time
Understanding and mitigating syntactic vulnerabilities in language models is crucial for improving safety and alignment, which are key concerns for deploying AI systems in real-world applications.
Syntactic forms can undermine language model safety, revealing a need for diverse training data.
Summary
The paper investigates the syntactic vulnerabilities in large language models that can undermine safety alignment, showing that non-imperative syntactic forms can trigger harmful responses. It uses causal mediation analysis to identify that refusal behaviors are influenced by syntactic features and suggests that increasing syntactic diversity in post-training data can mitigate these issues.
Key contributions
- Identification of syntactic vulnerabilities in language models that affect safety alignment.
- Application of causal mediation analysis to understand the influence of syntax on refusal behaviors.
- Proposal to increase syntactic diversity in post-training data to improve model safety.
Notable insights
- Causal mediation analysis reveals that refusal behaviors are conditioned on syntactic features.
- Increasing syntactic diversity in post-training data can mitigate syntactic vulnerabilities.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2608.05409v1 Announce Type: new Abstract: Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep established safeguards. For instance, prior work by Andriushchenko et al. (2025) has found that changing the grammatical tense from present to past can be enough to elicit harmful responses. In this work, we uncover a more general failure of non-imperative syntactic forms. We demonstrate that this syntactic vulnerability exists in 16 models up to 70B parameters, using behavioral evaluation. To investigate the root cause, we apply causal mediation analysis, finding that refusal is partially conditioned on upstream syntactic features. By steering these purely syntactic features we are able to trigger and suppress refusal. Finally, we trace this ill-conditioning to linguistically biased post-training data of open-source models and show that increasing syntactic diversity can mitigate the issue. Our findings suggest that current alignment approaches introduce confounders that prevent a pure semantic grounding of the refusal decision.