Skip to main content
QUICK REVIEW

[论文解读] The emojification of sentiment on social media: Collection and analysis of a longitudinal Twitter sentiment dataset

Wenjie Yin, Rabab Alkhalifa|arXiv (Cornell University)|Aug 31, 2021
Sentiment Analysis and Opinion Mining参考文献 14被引用 4
一句话总结

本论文提出 TM-Senti,一个大规模、弱监督的 Twitter 情感数据集,包含超过 1.84 亿条推文,覆盖七年时间,通过表情符号和表情图标进行标注。研究揭示了情感表达中表情符号使用量随时间增长的纵向趋势,且可通过互联网档案馆的存档实现数据集的完整重建,支持情感分析与文本分类研究的可复现性。

ABSTRACT

Social media, as a means for computer-mediated communication, has been extensively used to study the sentiment expressed by users around events or topics. There is however a gap in the longitudinal study of how sentiment evolved in social media over the years. To fill this gap, we develop TM-Senti, a new large-scale, distantly supervised Twitter sentiment dataset with over 184 million tweets and covering a time period of over seven years. We describe and assess our methodology to put together a large-scale, emoticon- and emoji-based labelled sentiment analysis dataset, along with an analysis of the resulting dataset. Our analysis highlights interesting temporal changes, among others in the increasing use of emojis over emoticons. We publicly release the dataset for further research in tasks including sentiment analysis and text classification of tweets. The dataset can be fully rehydrated including tweet metadata and without missing tweets thanks to the archive of tweets publicly available on the Internet Archive, which the dataset is based on.

研究动机与目标

  • 为解决社交媒体研究中缺乏纵向情感数据集的问题,特别是追踪情感随时间的演变。
  • 通过表情符号和表情图标作为情感标签的代理,构建大规模、弱监督的 Twitter 情感数据集。
  • 分析情感表达中的时间趋势,特别是用户生成内容中从表情符号向表情符号的转变。
  • 通过利用互联网档案馆公开存档的推文,确保数据集的可复现性与完整重建。
  • 通过开放数据发布,支持未来在情感分析、文本分类和社会媒体分析方面的研究。

提出的方法

  • 采用弱监督方法构建数据集,基于推文中是否存在表情符号(如 `:)`、`:(`)和表情图标(如 😊、😢)作为情感的代理标签。
  • 从互联网档案馆的 Twitter 数据集中收集了超过 1.84 亿条推文,时间跨度为 2013 年至 2020 年。
  • 根据推文中出现的表情符号和表情图标的感情极性,为每条推文分配情感标签,分类为正面、负面或中性。
  • 数据集包含完整的推文元数据(如用户 ID、时间戳、是否为转发),以支持纵向和上下文分析。
  • 通过验证流程确保标签一致性与数据质量,并在后续版本中进行了修正。
  • 可通过互联网档案馆的存档推文实现数据集的完全重建,确保可复现性。

实验结果

研究问题

  • RQ1从 2013 年到 2020 年,表情符号和表情图标在 Twitter 情感表达中的使用如何演变?
  • RQ2在社交媒体文本中,表情符号相较于传统表情符号,作为情感指标的可靠性如何?
  • RQ3在不同主题或事件中,情感分布的时间趋势是什么?
  • RQ4从表情符号向表情符号的转变如何反映数字交流行为的更广泛变化?
  • RQ5使用表情符号和表情图标进行弱监督的方法,能否生成可靠且可扩展的情感数据集,以支持纵向分析?

主要发现

  • 该数据集包含超过 1.84 亿条推文,情感标签基于表情符号和表情图标生成,支持大规模情感分析。
  • 在七年的期间内,表情符号的使用量显著且持续增长,表明情感表达正趋向“表情符号化”。
  • 包含表情符号的推文比例从 2013 年至 2020 年稳步上升,而表情符号的使用量则呈下降趋势,反映出数字表达方式的转变。
  • 可通过互联网档案馆的推文存档实现数据集的完全重建,确保可复现性并可访问原始推文元数据。
  • 弱监督标注方法生成了可靠且可扩展的数据集,适用于情感分析与文本分类的下游任务。
  • 在 v1 至 v3 版本中,数据集因附录中的拼写错误被修正,v3 版本于 2025 年 3 月发布。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。