A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard)
Xianling Zhang
Why It Matters
What makes this one worth your time
This work is relevant for AI researchers and engineers focused on improving safety and alignment in RL systems, particularly in contexts where data efficiency is crucial.
C-Guard enhances data-efficient RL alignment by optimizing conflicting training objectives.
Summary
The paper introduces C-Guard, a constitution-grid instrument designed to improve data efficiency in reinforcement learning (RL) alignment by addressing conflicting objectives in training, specifically focusing on the balance between catching real harm and avoiding over-refusal of benign prompts.
Key contributions
- Introduction of C-Guard for generating RL training data.
- Development of C-LIM for assessing learnability and optimizing data regions.
- Demonstration of improved learning impact metrics in specific data regions.
Notable insights
- C-LIM provides a per-cell learnability score to optimize data usage in RL training.
- The method identifies and prunes ineffective data regions before training, potentially saving resources.
Possible limitations
- Not stated in the abstract.
Abstract
arXiv:2608.00180v4 Announce Type: replace Abstract: Conflicting objectives are general in RL alignment, and training on them data-efficiently is hard. Training a safety guard with RL means optimizing two objectives that conflict: catch real harm, and do not refuse benign prompts. Our finding is that over-refusal improves 22.4% to 12.8%, while under-refusal on adversarial attacks silently worsens 0.27 to 0.33. We present C-Guard, a constitution-grid instrument that generates the RL training data, and C-LIM, a per-cell learnability score that decides each cell's move: prune, densify, amend, expand. C-LIM flags the dead-weight data region before any training budget is spent: 187 untargeted rows had bought zero gain, and our method lifts the same region's learning impact 0.733 to 0.80. Code and the constitution are open-sourced.