Skip to main content
QUICK REVIEW

[Paper Review] Harnessing Large Language Models' Empathetic Response Generation Capabilities for Online Mental Health Counselling Support

Siyuan Brandon Loh, Aravind Sesagiri Raamkumar|arXiv (Cornell University)|Oct 12, 2023
Mental Health via Writing4 citations
TL;DR

This study evaluates large language models (LLMs) and empathetic conversational systems (ECS) in generating empathetic responses for online mental health counselling. Using prompts from the EmpatheticDialogues dataset, LLMs outperformed both ECS models and human baselines in empathy metrics, particularly in exploring emotional themes beyond immediate input, suggesting LLMs are promising for scalable, empathetic mental health support with minimal fine-tuning.

ABSTRACT

Large Language Models (LLMs) have demonstrated remarkable performance across various information-seeking and reasoning tasks. These computational systems drive state-of-the-art dialogue systems, such as ChatGPT and Bard. They also carry substantial promise in meeting the growing demands of mental health care, albeit relatively unexplored. As such, this study sought to examine LLMs' capability to generate empathetic responses in conversations that emulate those in a mental health counselling setting. We selected five LLMs: version 3.5 and version 4 of the Generative Pre-training (GPT), Vicuna FastChat-T5, Pathways Language Model (PaLM) version 2, and Falcon-7B-Instruct. Based on a simple instructional prompt, these models responded to utterances derived from the EmpatheticDialogues (ED) dataset. Using three empathy-related metrics, we compared their responses to those from traditional response generation dialogue systems, which were fine-tuned on the ED dataset, along with human-generated responses. Notably, we discovered that responses from the LLMs were remarkably more empathetic in most scenarios. We position our findings in light of catapulting advancements in creating empathetic conversational systems.

Motivation & Objective

  • To assess the empathetic response generation capabilities of large language models (LLMs) in simulated online mental health counselling scenarios.
  • To compare LLMs against traditional empathetic conversational systems (ECS) and human-generated responses using standardized empathy metrics.
  • To investigate whether pre-trained LLMs can generate more empathetic responses than fine-tuned ECS models despite lacking task-specific fine-tuning.
  • To evaluate the impact of model architecture and prompt design on empathy in mental health-related dialogue generation.
  • To explore the potential of LLMs as scalable, low-resource alternatives to data-intensive ECS for mental health support.

Proposed method

  • Five LLMs were evaluated: GPT-3.5, GPT-4, Vicuna FastChat-T5, PaLM 2, and Falcon-7B-Instruct, all prompted with simple instructions to respond to utterances from the EmpatheticDialogues (ED) dataset.
  • Responses were evaluated using three automated empathy metrics: Interpretation (understanding emotional content), Emotional Reaction (empathetic tone), and Exploration (expansion beyond immediate input).
  • Baseline responses were drawn from the original ED dataset, representing human-generated, non-fine-tuned dialogue turns.
  • The evaluation framework compared LLMs, ECS models (fine-tuned on ED), and human responses across positive and negative sentiment contexts.
  • Statistical analysis assessed the significance of differences in empathy scores across model types and sentiment conditions.
  • A key methodological choice was using pre-trained LLMs with zero-shot prompting, avoiding fine-tuning to test their inductive bias for empathy.
Figure 1: Average proportion of responses in each model type with empathetic features. Scores are grouped by sentiment. Top panel: Emotional Reaction. Middle panel: interpretation. Bottom panel: Exploration
Figure 1: Average proportion of responses in each model type with empathetic features. Scores are grouped by sentiment. Top panel: Emotional Reaction. Middle panel: interpretation. Bottom panel: Exploration

Experimental results

Research questions

  • RQ1Can large language models generate more empathetic responses than fine-tuned empathetic conversational systems (ECS) in mental health counselling simulations?
  • RQ2How do LLMs compare to human-generated responses from the EmpatheticDialogues dataset in terms of empathy metrics like emotional understanding and response exploration?
  • RQ3Does the sentiment of the user’s input (positive vs. negative) affect the empathetic performance of LLMs and ECS models?
  • RQ4Is there a significant difference in empathy generation between LLMs and baseline models, particularly in exploring emotional themes beyond the immediate input?
  • RQ5To what extent can pre-trained LLMs achieve high empathy with minimal prompting, suggesting potential for low-data, scalable mental health chatbot deployment?

Key findings

  • LLMs significantly outperformed both ECS models and human baselines in generating empathetic responses, particularly in the Exploration metric, which measures response depth beyond the immediate input.
  • The interaction between LLMs and negative sentiment prompts was statistically significant (odds ratio: 1.32, p < 0.05), indicating LLMs were more likely to explore emotional themes in negative contexts.
  • LLMs showed stronger performance in the Interpretation metric for positive emotions, though their performance on negative emotions was less pronounced, suggesting a potential gap in handling distress.
  • Human-generated responses from the ED dataset performed worst overall, especially on Emotional Reaction and Exploration, highlighting dataset quality issues and limitations of the original ED data.
  • Despite no fine-tuning, LLMs achieved high empathy scores, suggesting that pre-trained models already encode rich empathetic reasoning from broad-scale pre-training.
  • The study found that GPT-3.5 and GPT-4 demonstrated superior empathy compared to other LLMs, though performance varied across models, indicating model-specific capabilities even within the same architecture family.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.