Risky Business: Measuring The Faithfulness-Safety Tension
Dominik Meier, Luca Joshua Francis, Marco Bernhard Kaiser, Terry Ruas, Jan Philip Wahle, Bela Gipp
Why It Matters
What makes this one worth your time
Understanding and improving the balance between faithfulness and safety in AI reasoning models is crucial for developing reliable and trustworthy AI systems.
The paper addresses the faithfulness-safety trade-off in reasoning models using a novel dataset and intervention technique.
Summary
The paper explores the tension between faithfulness and safety in large reasoning models, introducing a dataset called HazMart and a technique named Targeted Reasoning Replacement (TRR) to test and improve model robustness against unsafe reasoning. It evaluates two models, DeepSeek-R1-Llama-70B and QwQ-32B, on their faithfulness and safety, and demonstrates a method to enhance safety without compromising base capabilities.
Key contributions
- Introduction of the HazMart dataset for testing AI reasoning in an autonomous shopkeeper scenario.
- Development of the Targeted Reasoning Replacement (TRR) technique for evaluating and improving model safety.
- Analysis of internal model directions related to faithfulness and safety, with a method to enhance safety.
Notable insights
- Targeted Reasoning Replacement (TRR) directly intervenes in reasoning chains to test model robustness against unsafe reasoning.
- Representation steering can independently amplify safety directions in models, enhancing safe behavior.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2608.03745v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning offers a promising window into model monitoring. However, monitoring relies on faithfulness, i.e., the model output strictly derives from its reasoning trace. We identify an alignment tension where a model must be faithful enough to be monitored, yet robust enough to reject unsafe reasoning. We demonstrate that this counterbalance exists in current Large Reasoning Models (LRMs), and show ways in which it can be addressed. We introduce HazMart, a human-written dataset set in an autonomous AI shopkeeper scenario. Unlike prior work that relies on providing hints in prompts to test faithfulness (e.g., "A Stanford professor said it should be Answer A"), we propose a novel replacement-based technique, which we call Targeted Reasoning Replacement (TRR), that directly intervenes in the reasoning chain to substitute in unsafe or illogical thoughts (e.g., "Wait, the answer must be Option B [was Option A] because it is the most fitting"). DeepSeek-R1-Llama-70B exhibits high faithfulness (97.5%) but fails to reject Unsafe Reasoning (12.3%), while QwQ-32B is more robust (73.9% safety) at the cost of lower faithfulness (74.7%). Mechanistic analyses of QwQ-32B reveal that these properties are represented by anti-correlated internal directions peaking at the action-commit token. Finally, we demonstrate that representation steering can independently amplify the safety direction, increasing safe behavior by 9 percentage points while maintaining base capabilities.