Skip to main content
QUICK REVIEW

[论文解读] A Large-Scale Comparative Study of Accurate COVID-19 Information versus Misinformation

Yida Mu, Ye Jiang|arXiv (Cornell University)|Apr 10, 2023
Misinformation and Its Impacts被引用 4
一句话总结

本研究通过对2.42亿条与COVID-19相关的推文进行大规模对比分析,识别出准确信息与虚假信息之间的显著特征差异,采用了一项新创建的、基于证据的虚假信息分类数据集。结果表明,虚假信息的传播速度比准确内容快158%,与负面情绪和阴谋论主题高度相关,且被平台 disproportionately 地移除,尤其是与阴谋论相关的虚假信息;而与病毒起源相关的虚假信息则基本未受到处理。

ABSTRACT

The COVID-19 pandemic led to an infodemic where an overwhelming amount of COVID-19 related content was being disseminated at high velocity through social media. This made it challenging for citizens to differentiate between accurate and inaccurate information about COVID-19. This motivated us to carry out a comparative study of the characteristics of COVID-19 misinformation versus those of accurate COVID-19 information through a large-scale computational analysis of over 242 million tweets. The study makes comparisons alongside four key aspects: 1) the distribution of topics, 2) the live status of tweets, 3) language analysis and 4) the spreading power over time. An added contribution of this study is the creation of a COVID-19 misinformation classification dataset. Finally, we demonstrate that this new dataset helps improve misinformation classification by more than 9\% based on average F1 measure.

研究动机与目标

  • 理解社交媒体上准确与虚假COVID-19信息在统计和语言学上的差异。
  • 通过识别主题分布、语言使用和传播动态中的独特模式,应对疫情期间虚假信息传播的挑战。
  • 开发并验证一个新型高质量数据集,用于分类COVID-19虚假信息,以提升检测模型的性能。

提出的方法

  • 本研究通过Twitter API收集了超过2.42亿条与COVID-19相关的推文,使用经筛选的可信信息源列表和关键词。
  • 基于新创建的数据集,结合人工标注和后处理,训练了一种新型、基于证据的虚假信息分类器,以提高可靠性。
  • 采用主题建模和基于关键词的过滤方法,将虚假信息识别并分类为主题簇,包括阴谋论和对公共权威的批评等类别。
  • 使用词袋模型(Bag-of-Words)和LIWC(语言探究与词频统计)进行语言学分析,比较虚假信息与准确内容在情感基调、词频和语义类别上的差异。
  • 通过追踪转发和互动行为随时间的变化,测量信息的传播能力,包括持续时间与峰值传播速度。
  • 采用留案外交叉验证评估新分类器的性能,结果表明其F1分数显著优于先前的最先进模型。
A Large-Scale Comparative Study of Accurate COVID-19 Information versus Misinformation

实验结果

研究问题

  • RQ1准确信息与虚假信息在主题和语言使用上存在哪些差异?
  • RQ2社交媒体平台处理了哪些类型的虚假信息?与虚假信息实际传播情况相比如何?
  • RQ3不同类型虚假信息的传播能力随时间如何变化?与准确信息相比有何差异?

主要发现

  • 虚假信息的传播速度比准确信息快158%,其中与阴谋论相关的虚假信息传播持续时间最长,自发布后仍活跃超过32小时。
  • 与阴谋论相关的虚假信息占所有虚假信息来源的约三分之一,是平台审核的主要目标,超过40%的相关推文被移除。
  • 与病毒起源相关的推文获得的平台干预最少,发布后近70%仍可访问。
  • 使用LIWC的语言学分析显示,虚假信息与“愤怒”、“负面情绪”和“死亡”类别强烈相关,而准确信息则与“积极情绪”、“真实性”和“社会”类词汇相关。
  • 词袋模型分析识别出如“#nwoevilplans”、“#chinaliedandpeopledied”和“#plandemic”等关键词,高度指示虚假信息。
  • 新创建的数据集使虚假信息分类性能平均提升超过9%,F1分数达到0.70,优于基线模型的0.51。
Figure 1: (a) : Screenshot of an IFCN debunk post. The post includes 1) fact checking organisation, 2) misinformation claim, 3) explanation of why the claim is false and 4) the link to full debunk article. (b) : Partial screenshot of full debunk article of the IFCN debunk post.
Figure 1: (a) : Screenshot of an IFCN debunk post. The post includes 1) fact checking organisation, 2) misinformation claim, 3) explanation of why the claim is false and 4) the link to full debunk article. (b) : Partial screenshot of full debunk article of the IFCN debunk post.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。