[Paper Review] Empathic Conversations: A Multi-level Dataset of Contextualized Conversations
This paper introduces the Empathic Conversations dataset, consisting of 500 article-grounded dyadic dialogues with multi-level empathy and personality annotations, plus baseline models predicting empathy, distress, and related features.
Empathy is a cognitive and emotional reaction to an observed situation of others. Empathy has recently attracted interest because it has numerous applications in psychology and AI, but it is unclear how different forms of empathy (e.g., self-report vs counterpart other-report, concern vs. distress) interact with other affective phenomena or demographics like gender and age. To better understand this, we created the {\it Empathic Conversations} dataset of annotated negative, empathy-eliciting dialogues in which pairs of participants converse about news articles. People differ in their perception of the empathy of others. These differences are associated with certain characteristics such as personality and demographics. Hence, we collected detailed characterization of the participants' traits, their self-reported empathetic response to news articles, their conversational partner other-report, and turn-by-turn third-party assessments of the level of self-disclosure, emotion, and empathy expressed. This dataset is the first to present empathy in multiple forms along with personal distress, emotion, personality characteristics, and person-level demographic information. We present baseline models for predicting some of these features from conversations.
Motivation & Objective
- Collect 500 conversations between crowd workers discussing news articles with rich personality and demographic data.
- Capture self-reported empathy and distress, partner-perceived empathy, and turn-level annotations of empathy, emotion, and self-disclosure.
- Analyze correlations between empathy measures, demographics, and personality traits.
- Develop baseline models to predict turn-level empathy, emotion, polarity, and self-disclosure from text and context.
- Demonstrate the usefulness of multi-level annotations for modeling empathy in dialogue.
Proposed method
- Collect demographic and Big Five personality data via surveys before conversation collection.
- Ground conversations in one of 100 highly empathy-eliciting negative news articles; participants rate initial empathy/distress (Batson scale).
- Facilitate 15+ turn two-person conversations about the article; at end, participants rate their partner's perceived empathy.
- Obtain third-person turn-level annotations for Empathy, Self-Disclosure, Emotion, and Emotional Polarity; annotate a subset with dialogue acts.
- Train baseline models (Bi-RNN with attention and RoBERTa-base) to predict turn-level and conversation-level empathy-related labels; use 70/15/15 data split for turns and standard fine-tuning for essays.
- Explore correlations between personality/demographics and empathy, and assess simple textual feature correlations (e.g., pronoun usage) with perceived empathy.
Experimental results
Research questions
- RQ1What are the relationships between self-reported empathy/distress, partner-perceived empathy, and turn-level empathy annotations?
- RQ2How do demographic factors and personality traits relate to empathy and distress in conversations?
- RQ3How accurately can models predict turn-level empathy, emotion, polarity, and self-disclosure from text and context?
- RQ4How well can models predict perceived counterparty empathy across conversation participants?
- RQ5Do simple textual features relate to perceived empathy in dialogue?
Key findings
- 500 conversations with 5,821 turn-level annotations and 1,400 turn-level dialog acts were collected; dataset includes 79 participants with demographic and personality data.
- RoBERTa-base achieved the strongest turn-level predictions with Empathy 0.771, Emotion 0.814, Emotion Polarity 0.812, and Self-Disclosure 0.769 (Pearson r).
- Bi-RNN with attention and additional numeric features improved emotion, polarity, and self-disclosure predictions over the base Bi-RNN; RoBERTa-base generally outperformed neural baselines on turn-level tasks.
- Perceived counterparty empathy (Person 2) was more predictable (0.115) than for Person 1 (0.268) with Bi-RNN-att; essays-based empathy/distress predictions using RoBERTa-base yielded Empathy 0.560 and Distress 0.665 (mean 0.612).
- Pronoun usage in conversations showed the strongest correlation with perceived empathy among simple textual features (correlation 0.468).
- Gender, age, education showed various associations with self-reported and turn-level empathy/distress, indicating demographic and personality effects on empathetic expression.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.