Skip to main content
QUICK REVIEW

[論文レビュー] Empathic Conversations: A Multi-level Dataset of Contextualized Conversations

Damilola Omitaomu, Shabnam Tafreshi|arXiv (Cornell University)|May 25, 2022
Mental Health via Writing被引用数 24
ひとこと要約

この論文は Empathic Conversations データセットを紹介する。これは、複数レベルの共感と人格注釈を含む、記事を基盤とした500件の二者対話から成るデータセットおよび共感・苦痛、関連特徴を予測するベースラインモデルを含む。

ABSTRACT

Empathy is a cognitive and emotional reaction to an observed situation of others. Empathy has recently attracted interest because it has numerous applications in psychology and AI, but it is unclear how different forms of empathy (e.g., self-report vs counterpart other-report, concern vs. distress) interact with other affective phenomena or demographics like gender and age. To better understand this, we created the {\it Empathic Conversations} dataset of annotated negative, empathy-eliciting dialogues in which pairs of participants converse about news articles. People differ in their perception of the empathy of others. These differences are associated with certain characteristics such as personality and demographics. Hence, we collected detailed characterization of the participants' traits, their self-reported empathetic response to news articles, their conversational partner other-report, and turn-by-turn third-party assessments of the level of self-disclosure, emotion, and empathy expressed. This dataset is the first to present empathy in multiple forms along with personal distress, emotion, personality characteristics, and person-level demographic information. We present baseline models for predicting some of these features from conversations.

研究の動機と目的

  • 豊富な人格データおよび人口統計データを伴うニュース記事を議論するクラウドワーカー間の500件の対話を収集する。
  • 自己報告の共感と distress、パートナーが知覚する共感、および共感、感情、自己開示のターンレベル注釈を捉える。
  • 共感指標、人口統計、性格特性の相関を分析する。
  • テキストと文脈からターンレベルの共感、感情、極性、自己開示を予測するベースラインモデルを開発する。
  • 対話における共感のモデリングのための多層注釈の有用性を実証する。

提案手法

  • 会話収集前にアンケートを通じて人口統計データとBig Five性格データを収集する。
  • 会話は100件の高い共感喚起をもたらす負のニュース記事のいずれかに基づかせる;参加者は初期の共感/苦痛を評価する(Batsonスケール)。
  • 記事について15ターン以上の二人対話を実施する。終了時に参加者は相手の知覚された共感を評価する。
  • Empathy、Self-Disclosure、Emotion、Emotional Polarityのターンレベル注釈を第三者から取得する。対話行為を含むサブセットを注釈する。
  • ベースラインモデル(Bi-RNN with attention と RoBERTa-base)を訓練し、ターンレベルおよび会話レベルの共感関連ラベルを予測する。ターンには 70/15/15 のデータ分割を用い、エッセイには標準的なファインチューニングを適用する。
  • 人格/人口統計と共感の相関を探索し、簡単なテキスト特徴量(例: 代名詞使用)と知覚された共感の相関を評価する。

実験結果

リサーチクエスチョン

  • RQ1自己報告の共感/苦痛、パートナーが知覚する共感、ターンレベルの共感注釈の関係は何か。
  • RQ2人口統計的要因と性格特性が会話における共感と苦痛にどう関連するか。
  • RQ3テキストと文脈からターンレベルの共感、感情、極性、自己開示をモデルがどれだけ正確に予測できるか。
  • RQ4会話参加者間で知覚される相手の共感をモデルがどれくらい予測できるか。
  • RQ5簡単なテキスト特徴が対話における知覚された共感と関連するか。

主な発見

  • 500件の対話、5,821件のターンレベル注釈、および1,400件のターンレベルダイアログ行為を収集した。データセットには人口統計データと人格データを持つ79名の参加者が含まれる。
  • RoBERTa-baseはターンレベル予測で最も強く、Empathy 0.771、Emotion 0.814、Emotion Polarity 0.812、Self-Disclosure 0.769(Pearson r)を達成。
  • Bi-RNN with attentionと追加の数値特徴を用いたモデルは、基礎Bi-RNNより感情、極性、自己開示の予測を改善。RoBERTa-baseはターンレベルのタスクで一般にニューラルベースラインより優れていた。
  • 知覚された相手の共感(Person 2)は Bi-RNN-att で予測しやすく、0.115であったのに対し Person 1は0.268。RoBERTa-baseを用いたエッセイベースの共感/苦痛予測はEmpathy 0.560、Distress 0.665(平均0.612)を得た。
  • 会話における代名詞の使用は、単純なテキスト特徴の中で知覚された共感と最も強い相関を示し(相関 0.468)、最も高い。
  • 性別、年齢、教育は自己報告およびターンレベルの共感/苦痛とさまざまな関連を示し、人口統計と人格が共感表現に影響を及ぼすことを示唆している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。