Skip to main content
QUICK REVIEW

[Paper Review] Is ChatGPT More Empathetic than Humans?

Anuradha Welivita, Pearl Pu|arXiv (Cornell University)|Feb 22, 2024
Artificial Intelligence in Healthcare and Education14 citations
TL;DR

The study compares GPT-4–generated empathetic responses to human responses across 600 participants using a between-subjects design, finding GPT-4 often rated more empathetic, especially with an empathy-defined prompt.

ABSTRACT

This paper investigates the empathetic responding capabilities of ChatGPT, particularly its latest iteration, GPT-4, in comparison to human-generated responses to a wide range of emotional scenarios, both positive and negative. We employ a rigorous evaluation methodology, involving a between-groups study with 600 participants, to evaluate the level of empathy in responses generated by humans and ChatGPT. ChatGPT is prompted in two distinct ways: a standard approach and one explicitly detailing empathy's cognitive, affective, and compassionate counterparts. Our findings indicate that the average empathy rating of responses generated by ChatGPT exceeds those crafted by humans by approximately 10%. Additionally, instructing ChatGPT to incorporate a clear understanding of empathy in its responses makes the responses align approximately 5 times more closely with the expectations of individuals possessing a high degree of empathy, compared to human responses. The proposed evaluation framework serves as a scalable and adaptable framework to assess the empathetic capabilities of newer and updated versions of large language models, eliminating the need to replicate the current study's results in future research.

Motivation & Objective

  • Assess how empathetic GPT-4 (GPT-4) responses are in chitchat-style dialogues compared to human responses.
  • Evaluate two GPT-4 prompting strategies: vanilla (generic) and empathy-defined (cognitive, affective, and compassionate components).
  • Validate a scalable evaluation framework suitable for future LLM empathy assessments and generalize findings beyond a single model version.

Proposed method

  • Use EmpatheticDialogues dataset with 2,000 dialogues distributed across 32 emotions.
  • Conduct a between-groups study with 600 crowd workers evaluating responses from humans, GPT-4 (vanilla), and GPT-4 (empathy-defined).
  • Prompt GPT-4 with two instruction styles: vanilla and empathy-defined, to generate responses to the first turn of each dialogue.
  • Rate empathy on a three-point scale (Bad, Okay, Good) and analyze with one-way ANOVA and t-tests.
  • Measure evaluator empathy propensity using the Toronto Empathy Questionnaire (TEQ) and analyze its interaction with ratings.

Experimental results

Research questions

  • RQ1Does GPT-4 generate more empathic responses than humans in diverse emotional scenarios?
  • RQ2Does explicitly defining empathy in prompts improve GPT-4’s alignment with highly empathetic evaluators?
  • RQ3How do empathy ratings differ across positive versus negative emotional contexts?
  • RQ4Is there a relationship between raters’ inherent empathy (TEQ) and how they rate GPT-4 vs. human responses?

Key findings

  • GPT-4 (vanilla) and GPT-4 (empathy-defined) receive higher average empathy ratings than humans across all emotions.
  • GPT-4 (empathy-defined) yields the highest average ratings for all emotions and negative emotions, with about 11.21% and 9.61% increases respectively over humans.
  • GPT-4 (vanilla) shows a 13.14% higher average empathy rating for positive emotions than humans.
  • Differences between GPT-4 (empathy-defined) and GPT-4 (vanilla) are not statistically significant overall (p > 0.05).
  • Raters with higher empathy propensity tend to rate GPT-4 (empathy-defined) more highly, with a stronger slope than for humans or GPT-4 (vanilla).
  • Qualitative examples indicate GPT-4 can adopt non-directive, more empathetic communication when prompted with empathy-defined guidance.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.