Trust Stack for Mental Health AI: A Survey of Calibration across Human, Interaction, and AI Layers
Xin Sun, Yue Su, Yifan Mo, Qingyu Meng, Yuxuan Li, Min Chen, Mengyuan Zhang, Saku Sugawara, Charlotte Gerritsen, Sander L. Koole, Koen Hindriks, Jiahuan Pei
Why It Matters
What makes this one worth your time
This research addresses the critical need for trustworthy AI in mental health, which can significantly impact user outcomes and therapeutic practices.
A framework to align trust in AI for mental health support across diverse stakeholders.
Summary
The paper proposes a three-layer trust framework for AI systems in mental health support, integrating perspectives from various stakeholders and reviewing existing research on trust evaluation practices.
Key contributions
- Development of a three-layer trust framework for AI in mental health.
- Systematic review of existing AI-driven research and evaluation practices in the mental health domain.
- Outline of a research agenda for improving AI trustworthiness in mental health support.
Notable insights
- The distinction between human-oriented, AI-oriented, and interaction-oriented trust layers is a nuanced approach to understanding trust in AI systems.
- Identifying gaps between NLP metrics and real-world mental health needs highlights the importance of context in AI evaluation.
Possible limitations
- Not stated in the abstract.
Abstract
arXiv:2604.20166v3 Announce Type: replace Abstract: Language-based AI is increasingly deployed for mental health support, yet trust is evaluated in interdisciplinary but operationally misaligned ways: NLP and AI work measures robustness, safety, privacy, and explanations, while psychotherapy, HCI, and regulatory work emphasize therapeutic fidelity, lived experience, empathy, and reliance. Empathetic chatbots can elicit strong user trust without commensurate safety, while safer systems are under-trusted when their boundaries are opaque, a calibration gap no single community owns. Through a structured scoping synthesis of 61 papers, we survey this landscape into a three-layer framework separating (L1) human-oriented trust, (L2) interaction-oriented trustworthiness, and (L3) AI-oriented trustworthiness, and map five stakeholder perspectives onto these layers. We outline a research agenda for building socio-technically aligned trustworthy AI for mental health support, highlighting that the central objective should shift from maximizing perceived trust to calibrating human trust to demonstrated interaction- and AI-level trustworthiness.