Skip to main content
QUICK REVIEW

[論文レビュー] Exploring ChatGPT's Empathic Abilities

Kristina Schaaff, Caroline Reinig|arXiv (Cornell University)|Aug 7, 2023
Digital Mental Health Interventions被引用数 4
ひとこと要約

本研究では、標準化心理検査を用いて、ChatGPTの共感的機能を感情理解・表現、並列的感情反応、共感的パーソナリティの3つの次元から評価した。結果として、ChatGPTは91.7%のケースで感情を正しく特定し、70.7%の対話で並列的感情反応を示した。これはアスペルガーサンダ-症候群を有する個人より優れているが、健康な人間の平均より低い共感スコアを示した。

ABSTRACT

Empathy is often understood as the ability to share and understand another individual's state of mind or emotion. With the increasing use of chatbots in various domains, e.g., children seeking help with homework, individuals looking for medical advice, and people using the chatbot as a daily source of everyday companionship, the importance of empathy in human-computer interaction has become more apparent. Therefore, our study investigates the extent to which ChatGPT based on GPT-3.5 can exhibit empathetic responses and emotional expressions. We analyzed the following three aspects: (1) understanding and expressing emotions, (2) parallel emotional response, and (3) empathic personality. Thus, we not only evaluate ChatGPT on various empathy aspects and compare it with human behavior but also show a possible way to analyze the empathy of chatbots in general. Our results show, that in 91.7% of the cases, ChatGPT was able to correctly identify emotions and produces appropriate answers. In conversations, ChatGPT reacted with a parallel emotion in 70.7% of cases. The empathic capabilities of ChatGPT were evaluated using a set of five questionnaires covering different aspects of empathy. Even though the results show, that the scores of ChatGPT are still worse than the average of healthy humans, it scores better than people who have been diagnosed with Asperger syndrome / high-functioning autism.

研究の動機と目的

  • ChatGPTが人間らしく感情を理解し表現できるかどうかを調査すること。
  • 共感的感情の重要な側面である並列的感情反応の能力を評価すること。
  • 標準化心理検査を用いてChatGPTの共感的パーソナリティを評価すること。
  • ChatGPTの共感能力を、健康な個人およびアスペルガーサンダ-症候群を有する者を含む人間の基準と比較すること。
  • ChatGPTのような大規模言語モデルにおける共感の評価に再現可能であるフレームワークを確立すること。

提案手法

  • 本研究では、感情理解・表現、並列的感情反応分析、共感的パーソナリティ評価の3段階からなる方法論を採用した。
  • 感情理解の評価では、120件のプロンプトを用い、ChatGPTが特定の感情を表現するように文を再表現できるかを人間のアノテーションで評価した。
  • 並列的感情反応の評価では、120件の会話分析を通じて、ChatGPTがユーザーと同じ感情で応答する頻度を手動ラベリングで評価した。
  • 共感的パーソナリティの評価には、5つの検証済み心理検査(IRI、EQ、TEQ、PES、AQ)を用い、健康な男性・女性の母集団との基準データと比較した。
  • すべてのデータはEmpatheticDialoguesデータセットから抽出され、ChatGPTが生成した内容であり、人間のアノテーターが感情的内容をラベリングして検証を行った。
  • ChatGPTのスコアと人間の基準との間で統計的比較を行い、効果量およびパーセンテージ差を算出した。
Figure 2: Accuracy of Understanding and Expressing Emotions.
Figure 2: Accuracy of Understanding and Expressing Emotions.

実験結果

リサーチクエスチョン

  • RQ1ChatGPTは、ユーザーの入力に対して、意図された感情をどれくらいの割合で正しく識別・表現できるか。
  • RQ2ChatGPTは、ユーザーの感情状態を模倣する並列的感情反応をどれくらいの頻度で示すか。
  • RQ3ChatGPTの共感的パーソナリティスコアは、健康な人間およびアスペルガーサンダ-症候群を有する者と比較して、複数の共感次元でどのように異なるか。
  • RQ4標準化心理検査は、ChatGPTのような大規模言語モデルにおける共感を信頼性高く評価できるか。
  • RQ5部分的な共感能力を有するチャットボットを人間-コンピュータインタラクションに導入するにあたり、どのような倫理的含みがあるか。

主な発見

  • ChatGPTはテストケースの91.7%で意図された感情を正しく識別・表現し、感情理解および表現能力に優れていることが示された。
  • 会話の70.7%で、ChatGPTはユーザーの感情状態を模倣する並列的感情反応を示し、ユーザー感情の模倣能力が顕著に認められた。
  • ChatGPTの全体的な共感スコアは、健康な人間のスコアより顕著に低く、主な共感尺度では60%~77%低い効果量を示した。
  • 全体的な共感が低いにもかかわらず、ChatGPTはすべての共感検査でアスペルガーサンダ-症候群を有する者を上回り、特に認知的共感および感情調節の分野で顕著に優れていた。
  • ChatGPTは喜びを示す反応の傾向が強く、これはトレーニングデータのバイアスまたはプロンプト-応答パターンに起因する可能性がある。
  • 本研究は、言語モデルにおける共感の評価に有効な、複数の手法を統合した検証済みフレームワークを提供しており、今後の評価のテンプレートとして機能する。
Figure 4: Distribution of ChatGPT’s Emotional Responses.
Figure 4: Distribution of ChatGPT’s Emotional Responses.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。