Skip to main content
QUICK REVIEW

[论文解读] Characterizing COVID-19 Misinformation Communities Using a Novel Twitter Dataset

Shahan Ali Memon, Kathleen M. Carley|arXiv (Cornell University)|Aug 3, 2020
Misinformation and Its Impacts参考文献 33被引用 33
一句话总结

本文提出 CMU-MisCOV19,是一个大规模带注释的 Twitter 数据集,并在网络结构、社会语言学与错误信息传播参与度方面分析两类竞争性 COVID-19 错误信息社区(知情 vs 误导)。

ABSTRACT

From conspiracy theories to fake cures and fake treatments, COVID-19 has become a hot-bed for the spread of misinformation online. It is more important than ever to identify methods to debunk and correct false information online. In this paper, we present a methodology and analyses to characterize the two competing COVID-19 misinformation communities online: (i) misinformed users or users who are actively posting misinformation, and (ii) informed users or users who are actively spreading true information, or calling out misinformation. The goals of this study are two-fold: (i) collecting a diverse set of annotated COVID-19 Twitter dataset that can be used by the research community to conduct meaningful analysis; and (ii) characterizing the two target communities in terms of their network structure, linguistic patterns, and their membership in other communities. Our analyses show that COVID-19 misinformed communities are denser, and more organized than informed communities, with a possibility of a high volume of the misinformation being part of disinformation campaigns. Our analyses also suggest that a large majority of misinformed users may be anti-vaxxers. Finally, our sociolinguistic analyses suggest that COVID-19 informed users tend to use more narratives than misinformed users.

研究动机与目标

  • 创建一个多样化、带注释的 COVID-19 Twitter 数据集,并为错误信息分析建立一份全面的代码本。
  • 基于网络结构、语言模式和虚假信息归属,刻画知情与误导性社区。
  • 评估误导性群体中的机器人(Bots)参与和疫苗立场。
  • 提供数据与方法,促进可重复的错误信息检测研究。

提出的方法

  • 在三个数据收集日期使用多样化的 COVID-19 关键词和话题标签收集 Twitter 数据。
  • 将推文注释为 17 个类别,描述错误信息和真实信息类别。
  • 基于推文注释计算用户价值取向,以将用户划分为知情或误导性群体。
  • 用过滤为 COVID-19 相关推文的用户时间线扩充数据,用于网络、机器人与社会语言学分析。
  • 构建转发、提及和回复网络,并计算每个群体的网络密度。
  • 使用 Bot-Hunter 识别机器人,并使用双样本 z 检验比较两组之间的比例差异。
  • 对非机器人时间线执行基于 LIWC 的社会语言学分析,以比较两组的词汇类别。
  • 通过话题标签情感传播评估误导性用户中的疫苗立场。

实验结果

研究问题

  • RQ1是否能够创建一个多样化的带注释的 Twitter 数据集,用于 COVID-19 错误信息分析?
  • RQ2知情与误导性 COVID-19 社区在网络结构上有哪些差异?
  • RQ3误导性社群是否表现出比知情群体更高的机器人参与?
  • RQ4两组之间是否存在显著的社会语言学模式,如叙事用法或正式度?
  • RQ5误导性群体中的疫苗立场分布如何,机器人在其中扮演了怎样的角色?

主要发现

  • 误导性社区比知情社区更密集,表明回声室效应更强。
  • 约 47% 的用户为知情,29% 为误导性,24% 为模糊或无关。
  • 机器人占误导性用户的 19%,而知情用户为 11%,差异具有统计显著性。
  • 知情用户呈现更多叙事性语言、较高的代词和功能词使用以及更高的真实性。
  • 两组总体基调均为负面,误导性用户趋向使用更高比例的非正式语言(结果尚不确定)。
  • 在发表疫苗相关内容的误导性用户中,41% 是反疫苗者,22% 是支持疫苗者(37% 为模糊/不确定)。
  • 错误信息传播活动在误导性群体中更高的机器人存在率所提示。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。