Risk Governance for Generative AI Mental Health Support: A Multi-Turn Safety Architecture
Anabela C. Areias, Catarina Botelho, Ant\'onio Farinhas, Areti Vassilopoulos, Dora Janela, Xin Tong, Nuno M. Guerreiro, Maya D'Eon, Fab\'iola Costa, Ricardo Rei
Why It Matters
What makes this one worth your time
This work is crucial for deploying LLMs in sensitive areas like mental health, ensuring safer and more effective interactions.
A novel safety architecture enhances LLMs' handling of mental health risks in conversations.
Summary
The paper presents a model-agnostic safety governance architecture for large language models used in mental health support, focusing on contextual risk detection, reasoning-based verification, and protocol-guided response generation. The architecture was evaluated using synthetic conversations based on real-world narratives, showing high risk detection performance and improved clinician-preferred escalation responses.
Key contributions
- Development of a model-agnostic safety governance architecture for mental health support.
- Demonstration of improved risk detection and response quality in synthetic mental health conversations.
Notable insights
- Combining contextual risk detection with reasoning-based verification can improve response quality in mental health support.
- Protocol-guided response generation helps maintain rapport while managing risk.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2607.22692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk. Existing safety approaches primarily detect risk but rarely shape how models respond as conversational risk unfolds. We developed a model-agnostic safety governance architecture that combines contextual risk detection, reasoning-based verification, and protocol-guided response generation for multi-turn mental health interactions. Synthetic conversations grounded in real-world mental health narratives were used to evaluate the architecture's performance, tested with GPT-5-chat and Qwen3.5-27B, achieving high risk detection performance (specificity: 0.85 (95\%CI: 0.78;0.91), sensitivity: 0.92 (95\%CI: 0.88;0.95)) and increasing clinician-preferred escalation responses by 25.6--59.2pp while preserving rapport and connection. Performance remained stable across conversation length and generalized across both proprietary and open-source models. These findings demonstrate that clinically-grounded safety governance can extend beyond risk detection to improve how LLMs manage evolving mental health risk, providing a scalable framework for safer deployment across models.