Skip to main content
QUICK REVIEW

[论文解读] Discovering Basic Emotion Sets via Semantic Clustering on a Twitter Corpus

Eugene Yuta Bann|arXiv (Cornell University)|Dec 28, 2012
Sentiment Analysis and Opinion Mining参考文献 62被引用 3
一句话总结

本文提出一种数据驱动方法,通过在推特语料上进行语义聚类,发现基本情绪集合,采用潜在语义聚类(LSC)评估情绪术语的语义独特性。研究识别出一个新情绪集合——Accepting, Ashamed, Contempt, Interested, Joyful, Pleased, Sleepy, Stressed,其语义独特性相比伊萨克森的经典情绪集合提升了6.1%。

ABSTRACT

A plethora of words are used to describe the spectrum of human emotions, but how many emotions are there really, and how do they interact? Over the past few decades, several theories of emotion have been proposed, each based around the existence of a set of 'basic emotions', and each supported by an extensive variety of research including studies in facial expression, ethology, neurology and physiology. Here we present research based on a theory that people transmit their understanding of emotions through the language they use surrounding emotion keywords. Using a labelled corpus of over 21,000 tweets, six of the basic emotion sets proposed in existing literature were analysed using Latent Semantic Clustering (LSC), evaluating the distinctiveness of the semantic meaning attached to the emotional label. We hypothesise that the more distinct the language is used to express a certain emotion, then the more distinct the perception (including proprioception) of that emotion is, and thus more 'basic'. This allows us to select the dimensions best representing the entire spectrum of emotion. We find that Ekman's set, arguably the most frequently used for classifying emotions, is in fact the most semantically distinct overall. Next, taking all analysed (that is, previously proposed) emotion terms into account, we determine the optimal semantically irreducible basic emotion set using an iterative LSC algorithm. Our newly-derived set (Accepting, Ashamed, Contempt, Interested, Joyful, Pleased, Sleepy, Stressed) generates a 6.1% increase in distinctiveness over Ekman's set (Angry, Disgusted, Joyful, Sad, Scared). We also demonstrate how using LSC data can help visualise emotions. We introduce the concept of an Emotion Profile and briefly analyse compound emotions both visually and mathematically.

研究动机与目标

  • 通过分析社交媒体上情绪关键词的语言使用,识别语义最独特的的情绪集合。
  • 通过自然语言数据测量语义独特性,评估现有情绪理论的心理学有效性。
  • 开发一种基于真实语言数据迭代聚类的方法,以发现语义不可约的基本情绪集合。
  • 通过情绪轮廓和多维缩放技术可视化情绪状态,实现对复合情绪的分析。
  • 探索在临床心理学、情感工程和经济预测中通过实时情绪分析的应用。

提出的方法

  • 使用追踪的情绪关键词和过滤短语构建了一个包含超过21,000条推文的标注推特语料库。
  • 应用潜在语义聚类(LSC)通过向量表示的余弦相似度,分析情绪术语之间的语义相似性。
  • 使用部分奇异值分解(SVD)降低维度,并从情绪词共现矩阵中提取潜在语义结构。
  • 实施迭代LSC算法,通过最大化独特性,识别语义不可约的最优情绪集合。
  • 通过多维缩放技术开发“情绪轮廓”模型,以语义空间可视化情绪状态和复合情绪。
  • 通过地理相对性测试验证结果,并将LSC输出与伊萨克森和普鲁奇克等既定情绪模型进行比较。

实验结果

研究问题

  • RQ1在推特上的自然语言使用中,哪个现有情绪集合展现出最高的语义独特性?
  • RQ2能否通过真实语言表达的数据驱动聚类,发现一个语义上更具不可约性的新基本情绪集合?
  • RQ3如何利用情绪相关语言的语义聚类来可视化并数学建模复合情绪状态?
  • RQ4情绪术语的语义独特性在多大程度上与它们的心理独特性感知相关?
  • RQ5基于LSC的情绪分析能否检测出公众情绪表达中的偏见(如宣传偏见),相较于私人情绪语言?

主要发现

  • 在测试的六个情绪集合中,伊萨克森的经典情绪集合(Angry, Disgusted, Joyful, Sad, Scared)语义独特性最高。
  • 新推导出的情绪集合——Accepting, Ashamed, Contempt, Interested, Joyful, Pleased, Sleepy, Stressed——相比伊萨克森集合语义独特性提升了6.1%。
  • LSC算法成功识别出一个由八个语义不可约情绪组成的集合,更完整地代表了人类情绪体验的全谱。
  • 通过多维缩放生成的情绪轮廓能有效可视化个体和复合情绪状态,如“depressed”或“guilty”。
  • 分析表明,主要情绪的组合(如Joyful + Scared)与复合状态(如“depressed”)表现出可测量的相似性,支持对情绪混合的数学建模。
  • 本研究证明,社交媒体语言的语义聚类可检测情绪表达模式,具有在临床评估和经济预测中的潜在应用价值。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。