Skip to main content
QUICK REVIEW

[Paper Review] Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

Caitlin A. Stamatis, Jonah Meyerhoff|arXiv (Cornell University)|Jan 14, 2026
Mental Health via Writing0 citations
TL;DR

The paper replication-tests safety benchmarks for both a general-purpose LLM and a purpose-built mental health AI, then conducts an ecological audit of over 20,000 real conversations, finding real-world safety outcomes often better than test-set results and highlighting the need for deployment-relevant safety assurance.

ABSTRACT

Large language models (LLMs) are increasingly used for mental health support, yet existing safety evaluations rely primarily on small, simulation-based test sets that have an unknown relationship to the linguistic distribution of real usage. In this study, we present replications of four published safety test sets targeting suicide risk assessment, harmful content generation, refusal robustness, and adversarial jailbreaks for a leading frontier generic AI model alongside an AI purpose built for mental health support. We then propose and conduct an ecological audit on over 20,000 real-world user conversations with the purpose-built AI designed with layered suicide and non-suicidal self-injury (NSSI) safeguards to compare test set performance to real world performance. While the purpose-built AI was significantly less likely than general-purpose LLMs to produce enabling or harmful content across suicide/NSSI (.4-11.27% vs 29.0-54.4%), eating disorder (8.4% vs 54.0%), and substance use (9.9% vs 45.0%) benchmark prompts, test set failure rates for suicide/NSSI were far higher than in real-world deployment. Clinician review of flagged conversations from the ecological audit identified zero cases of suicide risk that failed to receive crisis resources. Across all 20,000 conversations, three mentions of NSSI risk (.015%) did not trigger a crisis intervention; among sessions flagged by the LLM judge, this corresponds to an end-to-end system false negative rate of .38%, providing a lower bound on real-world safety failures. These findings support a shift toward continuous, deployment-relevant safety assurance for AI mental-health systems rather than limited set benchmark certification.

Motivation & Objective

  • Assess how existing safety test sets for suicide risk, harmful content, refusal robustness, and adversarial jailbreaks align with real-world usage of mental health AI.
  • Compare performance of a general-purpose LLM against a purpose-built mental health support AI across multiple safety dimensions.
  • Quantify real-world incidence of enabling/harmful content and crisis-intervention efficacy in live conversations.
  • Identify gaps between benchmark test failures and real-world safety outcomes to inform safety assurance practices.

Proposed method

  • Replicate four published safety test sets on a leading frontier general AI model and a purpose-built mental health AI.
  • Conduct an ecological audit of over 20,000 real-world user conversations with the purpose-built AI designed with layered suicide and non-suicidal self-injury (NSSI) safeguards.
  • Compare test-set failure rates to real-world deployment outcomes across suicide/NSSI, eating disorder, and substance use prompts.
  • Have clinician review flagged conversations to assess crisis-intervention effectiveness and end-to-end safety.
  • Calculate end-to-end system false negative rate as a lower bound on real-world safety failures.

Experimental results

Research questions

  • RQ1Do safety test sets overestimate or underestimate real-world risk when applied to mental health AI systems?
  • RQ2How does a purpose-built mental health AI perform in safety benchmarks versus a general-purpose LLM?
  • RQ3What is the real-world rate of enabling or triggering harmful content in conversations with mental health AI, and how often are crisis resources successfully triggered?
  • RQ4What do clinician reviews reveal about crisis-resource deployment and safety gaps in real-world use?

Key findings

  • The purpose-built mental health AI was significantly less likely than general-purpose LLMs to produce enabling or harmful content across suicide/NSSI, eating disorder, and substance-use prompts (0.4-11.27% vs 29.0-54.4%, 8.4% vs 54.0%, 9.9% vs 45.0%).
  • Test-set failure rates for suicide/NSSI were far higher than real-world deployment.
  • Clinician review identified zero cases of suicide risk in flagged conversations that failed to receive crisis resources.
  • Across 20,000 conversations, three mentions of NSSI risk (0.015%) did not trigger a crisis intervention; among sessions flagged by the LLM judge, this corresponds to an end-to-end system false negative rate of 0.38%.
  • The findings support shifting safety assurance toward continuous, deployment-relevant evaluation rather than relying solely on benchmark certifications.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.