Skip to main content
QUICK REVIEW

[论文解读] Harnessing Large Language Models' Empathetic Response Generation Capabilities for Online Mental Health Counselling Support

Siyuan Brandon Loh, Aravind Sesagiri Raamkumar|arXiv (Cornell University)|Oct 12, 2023
Mental Health via Writing被引用 4
一句话总结

本研究评估了大型语言模型(LLMs)和共情对话系统(ECS)在生成在线心理健康咨询中的共情回应方面的能力。使用来自EmpatheticDialogues数据集的提示,LLMs在共情度量指标上优于ECS模型和人类基线,尤其是在超越即时输入的情感主题探索方面,表明LLMs在无需大量微调的情况下,具有作为可扩展、共情式心理健康支持的潜力。

ABSTRACT

Large Language Models (LLMs) have demonstrated remarkable performance across various information-seeking and reasoning tasks. These computational systems drive state-of-the-art dialogue systems, such as ChatGPT and Bard. They also carry substantial promise in meeting the growing demands of mental health care, albeit relatively unexplored. As such, this study sought to examine LLMs' capability to generate empathetic responses in conversations that emulate those in a mental health counselling setting. We selected five LLMs: version 3.5 and version 4 of the Generative Pre-training (GPT), Vicuna FastChat-T5, Pathways Language Model (PaLM) version 2, and Falcon-7B-Instruct. Based on a simple instructional prompt, these models responded to utterances derived from the EmpatheticDialogues (ED) dataset. Using three empathy-related metrics, we compared their responses to those from traditional response generation dialogue systems, which were fine-tuned on the ED dataset, along with human-generated responses. Notably, we discovered that responses from the LLMs were remarkably more empathetic in most scenarios. We position our findings in light of catapulting advancements in creating empathetic conversational systems.

研究动机与目标

  • 评估大型语言模型(LLMs)在模拟在线心理健康咨询场景中生成共情回应的能力。
  • 使用标准化的共情度量指标,将LLMs与传统的共情对话系统(ECS)及人工生成的回应进行比较。
  • 探究预训练LLMs是否能在缺乏任务特定微调的情况下,生成比微调过的ECS模型更具共情力的回应。
  • 评估模型架构与提示设计对心理健康相关对话生成中共情表现的影响。
  • 探索LLMs作为数据密集型ECS的可扩展、低资源替代方案在心理健康支持中的潜力。

提出的方法

  • 评估了五种LLMs:GPT-3.5、GPT-4、Vicuna FastChat-T5、PaLM 2和Falcon-7B-Instruct,均使用简单指令提示,响应EmpatheticDialogues(ED)数据集中的语句。
  • 使用三种自动化共情度量指标评估响应:理解(理解情感内容)、情感反应(共情语气)和探索(超越即时输入的扩展)。
  • 基线响应来自原始ED数据集,代表人工生成的、未经微调的对话回合。
  • 评估框架在积极和消极情绪语境下,比较LLMs、ECS模型(在ED上微调)和人工响应的表现。
  • 通过统计分析评估不同模型类型和情绪条件下的共情得分差异的显著性。
  • 一个关键的方法论选择是使用预训练LLMs进行零样本提示,避免微调,以检验其在共情方面的归纳偏置。
Figure 1: Average proportion of responses in each model type with empathetic features. Scores are grouped by sentiment. Top panel: Emotional Reaction. Middle panel: interpretation. Bottom panel: Exploration
Figure 1: Average proportion of responses in each model type with empathetic features. Scores are grouped by sentiment. Top panel: Emotional Reaction. Middle panel: interpretation. Bottom panel: Exploration

实验结果

研究问题

  • RQ1在心理健康咨询模拟中,大型语言模型能否生成比微调过的共情对话系统(ECS)更具共情力的回应?
  • RQ2LLMs在共情度量指标(如情感理解与回应探索)方面,与EmpatheticDialogues数据集中人工生成的回应相比如何?
  • RQ3用户输入的情绪(积极与消极)是否会影响LLMs和ECS模型的共情表现?
  • RQ4LLMs与基线模型之间在共情生成方面是否存在显著差异,特别是在超越即时输入的情感主题探索方面?
  • RQ5预训练LLMs在极少提示下能达到多高的共情水平,这是否表明其在低数据、可扩展的心理健康聊天机器人部署中具有潜力?

主要发现

  • LLMs在生成共情回应方面显著优于ECS模型和人类基线,尤其在衡量回应深度的‘探索’指标上表现突出。
  • LLMs与消极情绪提示之间的交互作用具有统计显著性(优势比:1.32,p < 0.05),表明LLMs在消极语境下更可能探索情感主题。
  • LLMs在积极情绪的理解指标上表现更强,尽管在消极情绪上的表现较弱,提示其在处理痛苦情绪方面可能存在差距。
  • 来自ED数据集的人工生成回应整体表现最差,尤其在情感反应和探索方面,凸显了数据集质量的问题以及原始ED数据的局限性。
  • 尽管未进行微调,LLMs仍取得了较高的共情得分,表明预训练模型已从大规模预训练中编码了丰富的共情推理能力。
  • 研究发现GPT-3.5和GPT-4在共情能力上优于其他LLMs,尽管各模型表现存在差异,表明即使在同一架构家族中,模型间也存在特定能力差异。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。