Skip to main content
QUICK REVIEW

[论文解读] Comparing the Efficacy of GPT-4 and Chat-GPT in Mental Health Care: A Blind Assessment of Large Language Models for Psychological Support

Birger Moëll|arXiv (Cornell University)|May 15, 2024
Artificial Intelligence in Healthcare and Education被引用 4
一句话总结

本研究开展了一项盲评、由临床医生主导的评估,比较了GPT-4与ChatGPT在回应18个与抑郁、焦虑、创伤及相关病症相关的心理提示时的表现。GPT-4显著优于ChatGPT,平均临床评分达到8.29/10,而ChatGPT为6.52/10,展现出更强的同理心、临床相关性及治疗指导能力。

ABSTRACT

Background: Rapid advancements in natural language processing have led to the development of large language models with the potential to revolutionize mental health care. These models have shown promise in assisting clinicians and providing support to individuals experiencing various psychological challenges. Objective: This study aims to compare the performance of two large language models, GPT-4 and Chat-GPT, in responding to a set of 18 psychological prompts, to assess their potential applicability in mental health care settings. Methods: A blind methodology was employed, with a clinical psychologist evaluating the models' responses without knowledge of their origins. The prompts encompassed a diverse range of mental health topics, including depression, anxiety, and trauma, to ensure a comprehensive assessment. Results: The results demonstrated a significant difference in performance between the two models (p > 0.05). GPT-4 achieved an average rating of 8.29 out of 10, while Chat-GPT received an average rating of 6.52. The clinical psychologist's evaluation suggested that GPT-4 was more effective at generating clinically relevant and empathetic responses, thereby providing better support and guidance to potential users. Conclusions: This study contributes to the growing body of literature on the applicability of large language models in mental health care settings. The findings underscore the importance of continued research and development in the field to optimize these models for clinical use. Further investigation is necessary to understand the specific factors underlying the performance differences between the two models and to explore their generalizability across various populations and mental health conditions.

研究动机与目标

  • 通过注册心理医生的盲评,评估GPT-4与ChatGPT在提供心理支持方面的临床有效性。
  • 比较两种大语言模型在涵盖抑郁、焦虑、创伤及情绪调节等多样化心理健康主题下的回应质量。
  • 评估大语言模型是否能够提供临床相关、富有同理心且目标导向的回应,适用于心理健康照护场景。
  • 通过识别性能差异与安全考量,为大语言模型在心理健康支持中的伦理整合提供建议。

提出的方法

  • 采用盲评设计,临床心理学家在不知晓模型来源的情况下对模型回应进行评分。
  • 使用了18个标准化的心理提示,涵盖焦虑、抑郁、自尊、压力及人际关系问题等关键领域。
  • 依据临床相关性、同理心、治疗准确性及目标导向性,对回应进行10分制评分。
  • 通过统计分析比较GPT-4与ChatGPT的平均回应得分,p < 0.05表示具有显著性差异。
  • 评估重点包括语言语气、情感反应能力及临床适宜性的定性与定量分析。
Figure 1 : Average rating of psychological advice generated by GPT-4 / Chat-GPT.
Figure 1 : Average rating of psychological advice generated by GPT-4 / Chat-GPT.

实验结果

研究问题

  • RQ1GPT-4与ChatGPT在回应心理提示时,其同理心与临床相关性表现如何比较?
  • RQ2哪一模型在应对抑郁、焦虑及创伤症状方面展现出更强的治疗一致性?
  • RQ3大语言模型在缺乏临床监督的情况下,能在多大程度上支持目标导向的心理干预?
  • RQ4两种模型在回应结构、情感基调及临床实用性方面存在哪些关键定性差异?

主要发现

  • GPT-4的平均临床评分为8.29/10,显著高于ChatGPT的6.52/10(p < 0.05)。
  • GPT-4展现出更强的同理心,其回应被认为更具同情心,并更敏锐地感知用户的情绪困扰。
  • GPT-4提供了更准确且可操作的临床策略,用于管理抑郁、焦虑及强迫症(OCD)等病症。
  • ChatGPT的回应被评价为缺乏足够细致性,且部分回应存在安全隐患,缺乏明确的治疗方向。
  • 临床心理学家判断,GPT-4通过语言与语气更有效地促进了治疗联盟的建立。
  • 性能差异在各类提示中保持一致,包括与创伤、注意力缺陷多动障碍(ADHD)及社交焦虑相关的提示。
Figure 2 : Rating for each response generated by GPT-4 / Chat-GPT.
Figure 2 : Rating for each response generated by GPT-4 / Chat-GPT.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。