[论文解读] Is ChatGPT More Empathetic than Humans?
本研究比较 GPT-4 生成的具同理心回应与人类回应,样本量为 600 名参与者,采用被试间设计,结果表明 GPT-4 在大多数情况下更具同理心,尤其是在以同理心为定义的提示下。
This paper investigates the empathetic responding capabilities of ChatGPT, particularly its latest iteration, GPT-4, in comparison to human-generated responses to a wide range of emotional scenarios, both positive and negative. We employ a rigorous evaluation methodology, involving a between-groups study with 600 participants, to evaluate the level of empathy in responses generated by humans and ChatGPT. ChatGPT is prompted in two distinct ways: a standard approach and one explicitly detailing empathy's cognitive, affective, and compassionate counterparts. Our findings indicate that the average empathy rating of responses generated by ChatGPT exceeds those crafted by humans by approximately 10%. Additionally, instructing ChatGPT to incorporate a clear understanding of empathy in its responses makes the responses align approximately 5 times more closely with the expectations of individuals possessing a high degree of empathy, compared to human responses. The proposed evaluation framework serves as a scalable and adaptable framework to assess the empathetic capabilities of newer and updated versions of large language models, eliminating the need to replicate the current study's results in future research.
研究动机与目标
- 评估 GPT-4 (GPT-4) 在闲聊风格对话中的回应有多大程度具有同理心,相较于人类回应。
- 评估两种 GPT-4 提示策略:vanilla(通用)和 empathy-defined(认知、情感和同情成分)。
- 验证一个可扩展的评估框架,适用于未来 LLM 同理心评估,并将发现推广到单一模型版本之外。
提出的方法
- 使用 EmpatheticDialogues 数据集,包含 2,000 组对话,分布在 32 种情感中。
- 进行一个被试间研究,600 名众包工作者评估人类、GPT-4(vanilla)和 GPT-4(empathy-defined)的回应。
- 用两种指令风格提示 GPT-4:vanilla 和 empathy-defined,以生成每段对话的第一轮回应。
- 在三点量表上对同理心进行评分(Bad、Okay、Good),并使用单因素方差分析和 t 检验进行分析。
- 使用 Toronto Empathy Questionnaire (TEQ) 评估评估者的同理心倾向,并分析其与评分的交互作用。
实验结果
研究问题
- RQ1GPT-4 在多样化情绪情境中是否比人类生成的回应更具同理心?
- RQ2在提示中明确定义同理心是否提高 GPT-4 与高度具同理心的评估者的一致性?
- RQ3积极情绪情境与消极情绪情境下的同理心评分有何差异?
- RQ4评估者本身的同理心(TEQ)与他们如何评估 GPT-4 与人类回应之间的关系是否存在关联?
主要发现
- GPT-4(vanilla)和 GPT-4(empathy-defined)在所有情感中的平均同理心评分均高于人类。
- GPT-4(empathy-defined)在所有情感和负面情感中获得最高的平均评分,分别比人类高约 11.21% 和 9.61%。
- GPT-4(vanilla)在正面情感上的平均同理心评分比人类高 13.14%。
- GPT-4(empathy-defined)与 GPT-4(vanilla)之间的差异总体上没有统计显著性(p > 0.05)。
- 具有更高同理心倾向的评估者往往更高地评价 GPT-4(empathy-defined),其斜率比对人类或 GPT-4(vanilla)更大。
- 定性示例表明,在同理心定义的指导下,GPT-4 可以采取非指令性、更加具同理心的沟通方式。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。