Skip to main content
QUICK REVIEW

[论文解读] Sentiment Expression via Emoticons on Social Media

Hao Wang, Jorge A. Castanon|arXiv (Cornell University)|Nov 9, 2015
Sentiment Analysis and Opinion Mining参考文献 17被引用 4
一句话总结

本文研究了社交媒体中表情符号在情感表达中的作用,表明虽然少数表情符号是强烈且可靠的积极或消极情感指标,但许多表情符号传达的是复杂或模糊的情感。通过四项实证分析——包括用户感知调查、聚类分析、有无表情符号的情感分类比较以及上下文分析——研究发现,表情符号显著影响情感分类性能,由于其意义具有细微差别且依赖于上下文,因此应谨慎使用。

ABSTRACT

Emoticons (e.g., :) and :( ) have been widely used in sentiment analysis and other NLP tasks as features to ma- chine learning algorithms or as entries of sentiment lexicons. In this paper, we argue that while emoticons are strong and common signals of sentiment expression on social media the relationship between emoticons and sentiment polarity are not always clear. Thus, any algorithm that deals with sentiment polarity should take emoticons into account but extreme cau- tion should be exercised in which emoticons to depend on. First, to demonstrate the prevalence of emoticons on social media, we analyzed the frequency of emoticons in a large re- cent Twitter data set. Then we carried out four analyses to examine the relationship between emoticons and sentiment polarity as well as the contexts in which emoticons are used. The first analysis surveyed a group of participants for their perceived sentiment polarity of the most frequent emoticons. The second analysis examined clustering of words and emoti- cons to better understand the meaning conveyed by the emoti- cons. The third analysis compared the sentiment polarity of microblog posts before and after emoticons were removed from the text. The last analysis tested the hypothesis that removing emoticons from text hurts sentiment classification by training two machine learning models with and without emoticons in the text respectively. The results confirms the arguments that: 1) a few emoticons are strong and reliable signals of sentiment polarity and one should take advantage of them in any senti- ment analysis; 2) a large group of the emoticons conveys com- plicated sentiment hence they should be treated with extreme caution.

研究动机与目标

  • 考察社交媒体文本中表情符号的普遍性及其情感信号作用。
  • 探究表情符号作为情感极性指标的可靠性和一致性。
  • 评估移除表情符号对情感分类性能的影响。
  • 理解在线交流中表情符号使用的语境和语义细微差别。
  • 为情感分析系统中合理谨慎地整合表情符号提供指导。

提出的方法

  • 分析大规模近期的Twitter数据集,以测量表情符号的频率和分布。
  • 开展用户感知调查,评估参与者对最常见表情符号的情感解读。
  • 对词语和表情符号进行聚类,以探索语义关联和上下文含义。
  • 比较微博帖子在移除表情符号前后的情感极性,以衡量情感变化。
  • 训练两个机器学习模型——一个包含表情符号,一个不包含——以评估其对分类准确率的影响。
  • 使用统计分析检验‘移除表情符号会降低情感分类性能’的假设。

实验结果

研究问题

  • RQ1表情符号在社交媒体文本中有多普遍,特别是在Twitter上?
  • RQ2表情符号在不同用户和语境下多大程度上能可靠地指示情感极性?
  • RQ3移除表情符号如何影响微博帖子的情感极性?
  • RQ4在情感分类任务中,包含表情符号是否能提升机器学习模型的性能?
  • RQ5表情符号使用的语义和语境细微差别是什么,这些特征为何会挑战直接的情感标注?

主要发现

  • 少数表情符号,如 :) 和 :(,分别是积极和消极情感的强而可靠的指标。
  • 大量表情符号传达复杂或模糊的情感,因此不适合直接分配情感极性。
  • 从文本中移除表情符号会显著改变微博帖子的感知情感,表明其具有重要的语义作用。
  • 使用表情符号进行训练的机器学习模型在性能上优于不使用表情符号的模型,证实了表情符号作为情感分类特征的价值。
  • 研究确认,应在情感分析中策略性地使用表情符号,对那些含义不明确或具有混合情感的符号应保持谨慎。
  • 用户对表情符号情感的感知存在差异,凸显了NLP系统中需采用上下文感知的解读方式。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。