Skip to main content
QUICK REVIEW

[论文解读] Global Reactions to COVID-19 on Twitter: A Labelled Dataset with Latent Topic, Sentiment and Emotion Attributes

Gupta Rk, Ajay Vishwanath|arXiv (Cornell University)|Feb 16, 2021
Misinformation and Its Impacts参考文献 17被引用 13
一句话总结

本文介绍了一个大规模、多属性标注的数据集,包含2020年1月至2021年1月期间1.32亿条Twitter推文,利用自然语言处理(NLP)和预训练模型,为每条推文标注了17个潜在属性:10个二元话题相关性标签、5个情绪强度评分(恐惧、愤怒、悲伤、喜悦、情感极性),以及2个定性情感与情绪类别。该数据集通过捕捉全球疫情相关情绪与情感的时序动态,为公共卫生、心理学和传播学等跨学科研究提供了支持。

ABSTRACT

This paper presents a large, labelled dataset on people's responses and expressions related to the COVID-19 pandemic over the Twitter platform. From 28 January 2020 to 1 Jan 2021, we retrieved over 132 million public Twitter posts (i.e., tweets) from more than 20 million unique users using four keywords: corona, wuhan, nCov and covid. Leveraging natural language processing techniques and pre-trained machine learning-based emotion analytic algorithms, we labelled each tweet with seventeen latent semantic attributes, including a) ten binary attributes indicating the tweet's relevance or irrelevance to the top ten detected topics, b) five quantitative emotion intensity attributes indicating the degree of intensity of the valence or sentiment (from extremely negative to extremely positive), and the degree of intensity of fear, of anger, of sadness and of joy emotions (from barely noticeable to extremely high intensity), and c) two qualitative attributes indicating the sentiment category and the dominant emotion category the tweet is mainly expressing. We report the descriptive statistics around the topic, sentiment and emotion attributes, and their temporal distributions, and discuss the dataset's possible usage in communication, psychology, public health, economics, and epidemiology research.

研究动机与目标

  • 创建一个全面的、大规模的数据集,以捕捉社交媒体上公众对2019冠状病毒病(COVID-19)大流行的感受与情绪反应。
  • 解决当前公开可用的、大规模多标注数据集缺乏的问题,这些数据集需整合话题相关性、情感强度与情绪类别。
  • 通过提供具有时间分辨率、语义丰富的推文标注,支持传播学、心理学、公共卫生与流行病学等领域的跨学科研究。
  • 利用预训练的自然语言处理(NLP)模型,自动化标注海量实时社交媒体数据中复杂的感情与语义属性。

提出的方法

  • 从2020年1月28日至2021年1月1日期间,使用四个关键词(corona、wuhan、nCov、covid)收集了1.32亿条公开的Twitter推文。
  • 应用自然语言处理(NLP)技术提取并分类语义话题,识别与十大疫情相关话题的相关性。
  • 采用基于预训练机器学习的情绪分析算法,量化五种核心情绪(恐惧、愤怒、悲伤、喜悦、情感极性)的强度,评分范围从几乎不可察觉到极高。
  • 为每条推文分配两个定性标签:一个用于主导情感类别(积极、消极、中性),一个用于主导情绪类别。
  • 采用结合主题建模与微调后的情绪分类模型的混合方法,确保数据集内的一致性与可扩展性。
  • 生成话题、情感与情绪属性的描述性统计与时间分布,以刻画随时间推移的公众话语动态。

实验结果

研究问题

  • RQ1公众对2019冠状病毒病(COVID-19)大流行的感受与情绪表达在不同话题之间以及随时间如何变化?
  • RQ2在全球疫情早期阶段,恐惧、愤怒、悲伤与喜悦在Twitter话语中的分布与强度如何?
  • RQ3情感与情绪属性如何与特定疫情相关话题(如封锁措施、疫苗或病例数)相关联?
  • RQ4预训练的NLP模型在实时社交媒体数据中大规模标注复杂感情与语义属性方面的准确性如何?
  • RQ5该多属性数据集在公共卫生、心理学与传播学等跨学科研究中可发挥何种作用?

主要发现

  • 该数据集包含来自超过2000万名独立用户的1.32亿条推文,收集时间覆盖疫情早期阶段的12个月。
  • 话题相关性标签表明,与病例数、政府应对措施及健康建议相关的推文是讨论最频繁的话题。
  • 情绪强度评分显示,恐惧与悲伤水平较高,尤其在病例数上升与封锁措施宣布期间。
  • 情感极性显示,随着时间推移,情绪从负面逐渐转向更中性,反映出公众话语的演变。
  • 主导情绪类别最常为“恐惧”,其次是“悲伤”与“愤怒”,表明公众对疫情相关不确定性的强烈情感反应。
  • 时间序列分析表明,情绪强度与话题相关性随全球事件(如疫情高峰与政策变化)而波动。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。