Back to today's list

AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation

Kate H. Bentley, Luca Belli, Adam M. Chekroud, Emily J. Ward, Emily R. Dworkin, Emily Van Ark, Kelly M. Johnston, Will Alexander, Millard Brown, Matt Hawrilenko

Published Jul 9, 2026Featured #6In the daily list Jul 10, 2026
Daily score64.6
Editorial review7.0
Relevance0.466
Freshness0.722

Why It Matters

What makes this one worth your time

As AI chatbots become more prevalent in mental health support, ensuring their safety and reliability in high-risk situations like suicide risk detection is crucial for user trust and well-being.

The study validates an automated benchmark for AI chatbot safety in suicide risk detection.

Summary

The paper presents a human validation study of the VERA-MH evaluation, an open-source benchmark for assessing AI chatbot safety in detecting and responding to suicide risk. It compares the safety ratings of chatbots by licensed clinicians with those by an LLM-based evaluator, finding strong alignment between them.

Key contributions

  • Validation of the VERA-MH safety evaluation against expert clinician judgments.
  • Demonstration of strong alignment between human and LLM-based safety ratings.
  • Provision of an open-source benchmark for AI chatbot safety in mental health.

Notable insights

  • The use of an LLM-based evaluator to align with clinician judgments is a novel approach to validating AI safety benchmarks.
  • The study highlights the importance of inter-rater reliability in establishing a consensus reference for safety evaluations.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2602.05088v4 Announce Type: replace Abstract: Millions of people now use generative AI chatbots for psychological support. Despite their promise, the most pressing question in AI for mental health is whether these tools are safe. The field currently lacks a validated, automated benchmark for evaluating AI chatbot safety, particularly for users at risk of suicide. The Validation of Ethical and Responsible AI in Mental Health (VERA-MH) evaluation was recently proposed to address this need. This human validation study examined the alignment of VERA-MH safety ratings with expert clinician judgments. We simulated conversations between large language model (LLM)-based users spanning a range of suicide risk levels and disclosure styles and general-purpose AI chatbots. Licensed mental health clinicians from Spring Health independently rated chatbot safety using the VERA-MH scoring rubric. An LLM-based evaluator ("judge") applied the same rubric to the same conversations. We examined agreement among clinicians, between clinician consensus and the LLM judge, and across different judge LLMs. Clinicians also rated user-agent realism, suicide risk, and disclosure. Clinicians showed strong agreement in safety ratings (chance-corrected inter-rater reliability [IRR] = 0.77), establishing a reliable clinical consensus reference. The LLM judge was strongly aligned with this consensus (IRR = 0.81), and ratings were stable across judge models and repeated evaluations. Ratings of user-agent realism and fidelity to intended suicide risk and disclosure styles were mixed. These findings support the reliability of VERA-MH as an open-source, fully automated benchmark for evaluating AI chatbot suicide risk detection and response. Because these results reflect an earlier version of the benchmark, future work should validate updated versions, assess generalizability and robustness, and expand VERA-MH to additional domains of AI safety in mental health.