Skip to main content
QUICK REVIEW

[论文解读] COVID-19 Twitter Dataset with Latent Topics, Sentiments and Emotions Attributes

Raj Kumar Gupta, Ajay Vishwanath|arXiv (Cornell University)|Jul 14, 2020
Misinformation and Its Impacts参考文献 23被引用 43
一句话总结

本论文提出一个大规模、具有全球代表性的Twitter数据集(2020年1月至2022年6月,共2.52亿条推文),包含细粒度的推文级标注,涵盖潜在主题、情感极性以及四种核心情绪(恐惧、愤怒、悲伤、快乐)。通过使用LDA进行主题建模,以及CrystalFeel模型进行情绪与情感评分,该数据集为公共卫生、心理学和社会科学等多学科研究提供了支持,能够实现对疫情相关话语的实时分析,具备丰富的语义意义属性。

ABSTRACT

This paper describes a large global dataset on people's discourse and responses to the COVID-19 pandemic over the Twitter platform. From 28 January 2020 to 1 June 2022, we collected and processed over 252 million Twitter posts from more than 29 million unique users using four keywords: "corona", "wuhan", "nCov" and "covid". Leveraging probabilistic topic modelling and pre-trained machine learning-based emotion recognition algorithms, we labelled each tweet with seventeen attributes, including a) ten binary attributes indicating the tweet's relevance (1) or irrelevance (0) to the top ten detected topics, b) five quantitative emotion attributes indicating the degree of intensity of the valence or sentiment (from 0: extremely negative to 1: extremely positive) and the degree of intensity of fear, anger, sadness and happiness emotions (from 0: not at all to 1: extremely intense), and c) two categorical attributes indicating the sentiment (very negative, negative, neutral or mixed, positive, very positive) and the dominant emotion (fear, anger, sadness, happiness, no specific emotion) the tweet is mainly expressing. We discuss the technical validity and report the descriptive statistics of these attributes, their temporal distribution, and geographic representation. The paper concludes with a discussion of the dataset's usage in communication, psychology, public health, economics, and epidemiology.

研究动机与目标

  • 开发一个全面的、具有全球代表性的、基于Twitter的新冠疫情公共话语数据集。
  • 解决在疫情相关社交媒体数据中,公开可用的、细粒度的、推文级别的主题、情感与情绪标注不足的问题。
  • 通过为每条推文提供丰富、具有语义与心理学意义的属性,支持多学科研究。
  • 支持对各国随时间推移的公众情感、情绪趋势与主题演变的纵向分析。

提出的方法

  • 使用Twitter标准API收集超过2.52亿条英文推文,关键词包括:'corona'、'wuhan'、'nCov'、'covid',以及后期的'vaccine'相关词汇。
  • 应用潜在狄利克雷分布(LDA)识别并为每条推文标注十个主要潜在主题。
  • 使用预训练的CrystalFeel模型预测恐惧、愤怒、悲伤与快乐情绪的情感极性(0–1)与强度得分(0–1)。
  • 基于定量得分生成分类情感(从极度负面到极度正面)与主导情绪(恐惧、愤怒、悲伤、快乐、无特定情绪)标签。
  • 将数据覆盖期扩展至2020年1月28日至2022年6月1日,包含从2021年11月起的8个月疫苗相关数据。
  • 确保符合Twitter的服务条款,并采用CC BY-NC 2.0许可,对商业用途施加限制。

实验结果

研究问题

  • RQ1在新冠疫情期间,不同国家的公众情感与情绪强度(恐惧、愤怒、悲伤、快乐)如何随时间演变?
  • RQ2Twitter上公众话语中的主导潜在主题是什么?它们如何响应公共卫生事件与政策变化而演变?
  • RQ3情感与情绪模式在多大程度上与流行病学趋势或政府干预措施相关?
  • RQ4不同国家用户的情感与情感特征有何差异?这些差异可能由哪些文化或语境因素解释?
  • RQ5推文中表达的情绪强度与类型在多大程度上可预测更广泛的公众心理健康趋势或行为反应?

主要发现

  • 该数据集包含2.52亿条推文,来自2900万名独立用户,时间范围为2020年1月28日至2022年6月1日,覆盖30个代表性国家。
  • LDA模型成功识别出十个主要潜在主题,每条推文均被赋予与每个主题簇的二值相关性得分。
  • 情感极性得分范围为0(极度负面)至1(极度正面),平均值显示情感倾向随时间从主要负面逐渐转向混合或正面。
  • 恐惧、愤怒、悲伤与快乐的情绪强度得分以0–1尺度量化,支持对情绪动态的纵向分析。
  • 数据集包含从2021年11月3日至2022年6月1日的8个月疫苗相关数据,支持对疫苗接种推广期间公众情感与情绪反应的分析。
  • 该数据集已在多项研究中被使用,其中一项研究显示,未来导向与更高的喜悦和愤怒相关,而过去导向则与恐惧和悲伤相关。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。