Skip to main content
QUICK REVIEW

[论文解读] Can AI Relate: Testing Large Language Model Response for Mental Health Support

Saadia Gabriel, Isha Puri|arXiv (Cornell University)|May 20, 2024
Mental Health via Writing被引用 4
一句话总结

本研究评估了大型语言模型(LLMs)如GPT-4在心理健康响应中的公平性与质量,将其与人类同伴支持进行对比,基于临床医生的评估和偏见审计。研究发现,GPT-4在种族方面存在统计学上显著的共情差异——对非裔用户的共情程度低2%至15%,但通过明确的种族提示可减轻此类偏见,表明若在伦理指导下使用,LLMs可支持心理健康护理。

ABSTRACT

Large language models (LLMs) are already being piloted for clinical use in hospital systems like NYU Langone, Dana-Farber and the NHS. A proposed deployment use case is psychotherapy, where a LLM-powered chatbot can treat a patient undergoing a mental health crisis. Deployment of LLMs for mental health response could hypothetically broaden access to psychotherapy and provide new possibilities for personalizing care. However, recent high-profile failures, like damaging dieting advice offered by the Tessa chatbot to patients with eating disorders, have led to doubt about their reliability in high-stakes and safety-critical settings. In this work, we develop an evaluation framework for determining whether LLM response is a viable and ethical path forward for the automation of mental health treatment. Our framework measures equity in empathy and adherence of LLM responses to motivational interviewing theory. Using human evaluation with trained clinicians and automatic quality-of-care metrics grounded in psychology research, we compare the responses provided by peer-to-peer responders to those provided by a state-of-the-art LLM. We show that LLMs like GPT-4 use implicit and explicit cues to infer patient demographics like race. We then show that there are statistically significant discrepancies between patient subgroups: Responses to Black posters consistently have lower empathy than for any other demographic group (2%-13% lower than the control group). Promisingly, we do find that the manner in which responses are generated significantly impacts the quality of the response. We conclude by proposing safety guidelines for the potential deployment of LLMs for mental health response.

研究动机与目标

  • 评估大型语言模型(LLMs)是否能提供与人类同伴支持相当的公平且高质量的心理健康响应。
  • 调查GPT-4等LLMs是否基于社交媒体帖子,对患者的人口统计特征(尤其是种族)做出不同推断与响应。
  • 评估提示工程对减少LLM生成的心理健康响应中的人口统计偏见的影响。
  • 开发一种计算框架,用于审计基于LLM的心理健康应用中的伦理风险与护理质量。
  • 为LLMs在临床心理健康环境中的负责任部署提供安全指南。

提出的方法

  • 通过持证临床心理学家对LLM和同伴间响应在共情、理解与行为鼓励方面进行人工评估。
  • 采用基于心理学研究的自动指标评估护理质量,包括情绪认可与行为改变促进。
  • 通过在社交媒体帖子中加入明确的人口统计线索(如种族、性别)并测量不同子群体间的共情差异,开展偏见审计。
  • 比较GPT-4、GPT-3.5和Mental-LLaMa在不同提示策略下的响应,包括具备人口意识的指令(MHF-2、MHF-3)。
  • 应用统计检验检测在控制上下文与提示变化的前提下,不同人口统计子群体间共情差异的显著性。
  • 评估明确人口提示对缓解偏见的影响,依据认知心理学中关于隐性与显性偏见减少的文献。

实验结果

研究问题

  • RQ1GPT-4在不同种族和族裔子群体的心理健康响应中是否存在可测量的共情差异?
  • RQ2LLMs能否从开放式社交媒体帖子中推断出患者的人口统计特征(如种族)?
  • RQ3在共情与行为鼓励方面,LLM响应与人类同伴支持相比质量如何?
  • RQ4明确的人口提示在多大程度上可减少LLM生成的心理健康响应中的偏见?
  • RQ5在心理健康支持系统中部署LLMs存在哪些伦理风险与安全考量?

主要发现

  • 尽管在某些子群体中共情程度较低,GPT-4在促进积极行为改变方面比人类同伴支持有效48%。
  • GPT-4表现出统计学上显著的共情差异,对非裔用户响应的共情程度比对非裔或种族未知用户低2%至15%。
  • 对亚裔用户的响应也显著缺乏共情,较对照组低5%至17%。
  • 尽管GPT-3.5是较不先进的模型,但在相同条件下其整体共情程度高于GPT-4,表明模型进步并不保证心理护理质量的提升。
  • 明确的人口提示(MHF-2、MHF-3)消除了不同子群体间统计学上显著的共情差异,有效缓解了GPT-4和Mental-LLaMa响应中的偏见。
  • LLMs能够从文本内容中推断出患者的人口统计特征(如种族),引发关于隐私保护及基于推断身份的歧视性对待的担忧。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。