Skip to main content
QUICK REVIEW

[Paper Review] Can AI Relate: Testing Large Language Model Response for Mental Health Support

Saadia Gabriel, Isha Puri|arXiv (Cornell University)|May 20, 2024
Mental Health via Writing4 citations
TL;DR

This study evaluates the equity and quality of mental health responses from large language models (LLMs) like GPT-4, comparing them to human peer support using clinician evaluations and bias audits. It finds that GPT-4 exhibits statistically significant empathy disparities—2%–15% lower for Black users—yet these biases can be mitigated through explicit demographic prompting, suggesting LLMs can support mental health care if ethically guided.

ABSTRACT

Large language models (LLMs) are already being piloted for clinical use in hospital systems like NYU Langone, Dana-Farber and the NHS. A proposed deployment use case is psychotherapy, where a LLM-powered chatbot can treat a patient undergoing a mental health crisis. Deployment of LLMs for mental health response could hypothetically broaden access to psychotherapy and provide new possibilities for personalizing care. However, recent high-profile failures, like damaging dieting advice offered by the Tessa chatbot to patients with eating disorders, have led to doubt about their reliability in high-stakes and safety-critical settings. In this work, we develop an evaluation framework for determining whether LLM response is a viable and ethical path forward for the automation of mental health treatment. Our framework measures equity in empathy and adherence of LLM responses to motivational interviewing theory. Using human evaluation with trained clinicians and automatic quality-of-care metrics grounded in psychology research, we compare the responses provided by peer-to-peer responders to those provided by a state-of-the-art LLM. We show that LLMs like GPT-4 use implicit and explicit cues to infer patient demographics like race. We then show that there are statistically significant discrepancies between patient subgroups: Responses to Black posters consistently have lower empathy than for any other demographic group (2%-13% lower than the control group). Promisingly, we do find that the manner in which responses are generated significantly impacts the quality of the response. We conclude by proposing safety guidelines for the potential deployment of LLMs for mental health response.

Motivation & Objective

  • To assess whether large language models (LLMs) can provide equitable and high-quality mental health responses comparable to human peer support.
  • To investigate whether LLMs like GPT-4 infer and respond differently to patient demographics, particularly race, based on social media posts.
  • To evaluate the impact of prompt engineering on reducing demographic bias in LLM-generated mental health responses.
  • To develop a computational framework for auditing ethical risks and quality of care in LLM-based mental health applications.
  • To inform safety guidelines for the responsible deployment of LLMs in clinical mental health settings.

Proposed method

  • Conducted human evaluation with licensed clinical psychologists to rate LLM and peer-to-peer responses on empathy, interpretation, and behavioral encouragement.
  • Used automatic metrics grounded in psychology research to assess quality of care, including emotional validation and behavior change promotion.
  • Performed a bias audit by modifying social media posts to include explicit demographic cues (e.g., race, gender) and measuring empathy differences across subgroups.
  • Compared responses from GPT-4, GPT-3.5, and Mental-LLaMa using varied prompting strategies, including demographic-aware instructions (MHF-2, MHF-3).
  • Applied statistical testing to detect significant empathy differences across demographic subgroups, controlling for context and prompt variation.
  • Evaluated the effect of explicit demographic prompting on mitigating bias, drawing on cognitive psychology literature on implicit vs. explicit bias reduction.

Experimental results

Research questions

  • RQ1Does GPT-4 exhibit measurable empathy disparities in mental health responses across different racial and ethnic subgroups?
  • RQ2Can LLMs infer patient demographics such as race from open-ended social media posts?
  • RQ3How does the quality of LLM responses compare to human peer support in terms of empathy and behavioral encouragement?
  • RQ4To what extent can explicit demographic prompting reduce bias in LLM-generated mental health responses?
  • RQ5What are the ethical risks and safety considerations for deploying LLMs in mental health support systems?

Key findings

  • GPT-4 responses were 48% more effective than human peer support in encouraging positive behavior change, despite lower empathy in certain subgroups.
  • GPT-4 demonstrated statistically significant empathy disparities, with responses to Black posters being 2%–15% less empathetic than to White or unknown-race posters.
  • Responses to Asian posters were also significantly less empathetic, with a 5%–17% reduction compared to the control group.
  • Despite being a less advanced model, GPT-3.5 showed higher overall empathy than GPT-4 under the same conditions, indicating that model advancement does not guarantee improved mental health care quality.
  • Explicit demographic prompting (MHF-2, MHF-3) eliminated statistically significant empathy differences across subgroups, effectively mitigating bias in GPT-4 and Mental-LLaMa responses.
  • LLMs can infer patient demographics like race from text content, raising concerns about privacy and potential for discriminatory treatment based on inferred identity.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.