[论文解读] Exploring ChatGPT's Empathic Abilities
本研究通过标准化心理问卷,从情绪理解与表达、平行情感反应、共情人格三个维度评估ChatGPT的共情能力。结果显示,ChatGPT在91.7%的案例中正确识别情绪,并在70.7%的互动中以平行情绪作出回应,其表现优于阿斯伯格综合征患者,但在共情指标上仍低于健康人类平均水平。
Empathy is often understood as the ability to share and understand another individual's state of mind or emotion. With the increasing use of chatbots in various domains, e.g., children seeking help with homework, individuals looking for medical advice, and people using the chatbot as a daily source of everyday companionship, the importance of empathy in human-computer interaction has become more apparent. Therefore, our study investigates the extent to which ChatGPT based on GPT-3.5 can exhibit empathetic responses and emotional expressions. We analyzed the following three aspects: (1) understanding and expressing emotions, (2) parallel emotional response, and (3) empathic personality. Thus, we not only evaluate ChatGPT on various empathy aspects and compare it with human behavior but also show a possible way to analyze the empathy of chatbots in general. Our results show, that in 91.7% of the cases, ChatGPT was able to correctly identify emotions and produces appropriate answers. In conversations, ChatGPT reacted with a parallel emotion in 70.7% of cases. The empathic capabilities of ChatGPT were evaluated using a set of five questionnaires covering different aspects of empathy. Even though the results show, that the scores of ChatGPT are still worse than the average of healthy humans, it scores better than people who have been diagnosed with Asperger syndrome / high-functioning autism.
研究动机与目标
- 探究ChatGPT是否能够以类人方式理解并表达情绪。
- 评估ChatGPT在平行情感反应方面的能力,这是情感共情的关键组成部分。
- 使用标准化心理问卷评估ChatGPT的共情人格特质。
- 将ChatGPT的共情表现与人类基准进行比较,包括健康个体和阿斯伯格综合征患者。
- 建立可复现的框架,用于评估类似ChatGPT的大型语言模型的共情能力。
提出的方法
- 本研究采用三部分方法:情绪理解与表达、平行情感反应分析、共情人格评估。
- 在情绪理解方面,使用120个提示测试ChatGPT是否能重新表述语句以表达特定情绪,准确率通过人工标注进行衡量。
- 在平行情感反应方面,分析120段对话,判断ChatGPT是否以与用户相同的情绪作出回应,采用人工标注。
- 共情人格通过五种经过验证的心理问卷进行评估:IRI、EQ、TEQ、PES和AQ,得分与健康男性和女性的常模数据进行比较。
- 所有数据均来自EmpatheticDialogues数据集,由ChatGPT生成,人工标注者对情感内容进行标注以验证。
- 对ChatGPT的得分与人类基准进行统计比较,包括效应量和百分比差异。

实验结果
研究问题
- RQ1在多大程度上,ChatGPT能够正确识别并表达用户输入中的特定情绪?
- RQ2ChatGPT在多大频率上产生与用户情绪状态一致的平行情感反应?
- RQ3在多个共情维度上,ChatGPT的共情人格得分与健康人类及阿斯伯格综合征患者相比如何?
- RQ4标准化心理问卷是否能够可靠地评估类似ChatGPT的大型语言模型的共情能力?
- RQ5在人机交互中部署具备部分共情能力的聊天机器人,其伦理影响是什么?
主要发现
- 在91.7%的测试案例中,ChatGPT正确识别并表达了预期情绪,显示出强大的情绪理解与表达能力。
- 在70.7%的对话中,ChatGPT以平行情绪状态作出回应,表明其具有显著的模仿用户情绪的能力。
- ChatGPT的整体共情得分显著低于健康人类,关键共情量表的效应量差异在60%至77%之间。
- 尽管整体共情水平较低,但ChatGPT在所有共情问卷中均优于阿斯伯格综合征患者,尤其在认知共情和情绪调节方面表现更优。
- ChatGPT表现出强烈倾向以喜悦回应,这可能反映了训练数据中的偏差或提示-响应模式。
- 本研究提供了一个经过验证的多方法共情评估框架,为未来语言模型的共情评估提供了模板。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。