Skip to main content
QUICK REVIEW

[論文レビュー] Discovering Basic Emotion Sets via Semantic Clustering on a Twitter Corpus

Eugene Yuta Bann|arXiv (Cornell University)|Dec 28, 2012
Sentiment Analysis and Opinion Mining参考文献 62被引用数 3
ひとこと要約

本稿では、Twitterコーパス上での意味的クラスタリングを用いて、データ駆動型のアプローチで基本感情セットを発見する手法を提案している。Latent Semantic Clustering (LSC) を用いて、感情語の意味的明確性を評価した。本研究では、エクマンの古典的セットよりも意味的により明確な新しいセットである「受容的、恥ずかしい、軽蔑、関心、喜び、満足、眠い、ストレス」の8つの感情を同定し、エクマンのセット比で6.1%の明確性向上を達成した。

ABSTRACT

A plethora of words are used to describe the spectrum of human emotions, but how many emotions are there really, and how do they interact? Over the past few decades, several theories of emotion have been proposed, each based around the existence of a set of 'basic emotions', and each supported by an extensive variety of research including studies in facial expression, ethology, neurology and physiology. Here we present research based on a theory that people transmit their understanding of emotions through the language they use surrounding emotion keywords. Using a labelled corpus of over 21,000 tweets, six of the basic emotion sets proposed in existing literature were analysed using Latent Semantic Clustering (LSC), evaluating the distinctiveness of the semantic meaning attached to the emotional label. We hypothesise that the more distinct the language is used to express a certain emotion, then the more distinct the perception (including proprioception) of that emotion is, and thus more 'basic'. This allows us to select the dimensions best representing the entire spectrum of emotion. We find that Ekman's set, arguably the most frequently used for classifying emotions, is in fact the most semantically distinct overall. Next, taking all analysed (that is, previously proposed) emotion terms into account, we determine the optimal semantically irreducible basic emotion set using an iterative LSC algorithm. Our newly-derived set (Accepting, Ashamed, Contempt, Interested, Joyful, Pleased, Sleepy, Stressed) generates a 6.1% increase in distinctiveness over Ekman's set (Angry, Disgusted, Joyful, Sad, Scared). We also demonstrate how using LSC data can help visualise emotions. We introduce the concept of an Emotion Profile and briefly analyse compound emotions both visually and mathematically.

研究の動機と目的

  • SNS上での感情キーワード周辺の言語使用を分析することで、意味的に最も明確な感情セットを特定すること。
  • 自然言語データを用いた意味的明確性の測定により、既存の感情理論の心理的妥当性を評価すること。
  • 現実世界の言語的データに対する反復的クラスタリングを用いて、意味的に不可削減な基本感情セットを発見する手法を開発すること。
  • 感情プロファイルと多次元尺度法を用いて感情状態を可視化し、複合感情の分析を可能にすること。
  • 臨床心理学、感情工学、経済予測への応用を検討し、リアルタイムでの感情分析を可能にすること。

提案手法

  • トラッキングされた感情キーワードとフィルタリングされたフレーズを用いて、21,000件を超えるラベル付きTwitterコーパスを構築した。
  • 感情語間の意味的類似性を分析するため、ベクトル表現のコサイン類似度を用いてLatent Semantic Clustering (LSC) を適用した。
  • 感情語の共起行列からの潜在的意味構造を抽出し、次元削減のために部分的特異値分解 (SVD) を用いた。
  • 意味的明確性を最大化することで、意味的に不可削減な感情セットを特定する反復的LSCアルゴリズムを実装した。
  • 多次元尺度法を用いて感情プロファイルモデルを開発し、意味空間上での感情状態および複合感情の可視化を可能にした。
  • 地理的相関テストを用いて結果を検証し、LSCの出力結果をエクマンやプラッチクのモデルと比較した。

実験結果

リサーチクエスチョン

  • RQ1Twitter上での自然言語使用において、既存のどの感情セットが最も意味的に明確であるか?
  • RQ2現実世界の言語的表現のデータ駆動的クラスタリングにより、意味的により不可削減な新しい基本感情セットを発見できるか?
  • RQ3感情関連言語の意味的クラスタリングをどのようにして、複合感情状態の可視化および数学的モデル化に活用できるか?
  • RQ4感情語の意味的明確性と、それらが人々にどのように心理的に明確に感じられるかとの相関はどの程度強いのか?
  • RQ5LSCに基づく感情分析は、公的表現におけるバイアス(例:広報バイアス)を、私的表現と比較して検出できるか?

主な発見

  • 6つのテストされた感情セットの中で、エクマンの古典的セット(怒り、嫌悪、喜び、悲しみ、恐れ)が最も意味的に明確であることが判明した。
  • 新たに導出された感情セット「受容的、恥ずかしい、軽蔑、関心、喜び、満足、眠い、ストレス」は、エクマンのセット比で6.1%の意味的明確性向上を達成した。
  • LSCアルゴリズムは、人間の感情経験の全範囲をよりよく反映する8つの感情の意味的に不可削減なセットを的確に同定できた。
  • 多次元尺度法を用いて生成された感情プロファイルは、個々の感情状態および複合感情状態(例:「うつ」や「罪悪感」)を効果的に可視化できた。
  • 主な感情の組み合わせ(例:喜び+恐れ)が、「うつ」のような複合状態と測定可能な類似性を示したため、感情の混合状態の数学的モデル化が有効であることが裏付けられた。
  • 本研究では、SNS言語の意味的クラスタリングが、臨床的評価や経済予測への応用が可能な感情表現のパターンを検出できることを示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。