Cat-DPO: Category-Adaptive Safety Alignment
Tiankai Yang, Yi Nian, Xinyuan Li, Ruiyao Xu, Henry Peng Zou, Kaize Ding, Xiyang Hu, Yan Liu, Yue Zhao
Why It Matters
What makes this one worth your time
This work is significant for AI researchers and engineers focused on enhancing the safety and reliability of language models, particularly in addressing the nuanced challenges of harmful responses.
Cat-DPO offers a novel approach to safety alignment in language models by adapting safety margins per harm category.
Summary
The paper presents Cat-DPO, a direct-preference-optimization algorithm that addresses safety alignment in large language models by implementing a category-adaptive safety margin, improving both helpfulness and harmlessness across different harm categories.
Key contributions
- Introduction of Cat-DPO, a category-adaptive safety alignment algorithm.
- Demonstration of improved aggregate helpfulness and harmlessness in language models.
- Reduction of per-category safety variance and the best-to-worst gap in safety performance.
Notable insights
- The adaptive safety margin allows for dynamic adjustments based on the model's performance in specific harm categories, which could lead to more effective safety measures.
- The approach contrasts with traditional methods that apply a uniform safety measure, potentially leading to overlooked vulnerabilities in certain categories.
Possible limitations
- Not stated in the abstract.
Abstract
arXiv:2604.17299v3 Announce Type: replace-cross Abstract: Aligning large language models with human preferences must balance two competing goals: responding helpfully to legitimate requests and reliably refusing harmful ones. Most preference-based safety alignment methods collapse safety into a single scalar that is applied uniformly to every preference pair. The result is a model that looks safe on average but stays relatively unsafe on a minority of harm categories. We cast safety alignment as a per-category constrained optimization problem and derive Cat-DPO, a direct-preference-optimization algorithm with a separate adaptive safety margin for each harm category. The margin tightens when the model still produces unsafe responses on a category and relaxes once the model catches up, so the training signal tracks each category's current difficulty rather than averaging under one global rate. Across two LLM backbones and six preference-learning baselines, Cat-DPO improves aggregate helpfulness and harmlessness and compresses per-category safety variance and the best-to-worst gap, offering a drop-in per-category refinement of direct preference safety alignment.