[论文解读] CMU-MisCov19: A Novel Twitter Dataset for Characterizing COVID-19 Misinformation
本文介绍了CMU-MisCov19,一个大规模、人工标注的Twitter数据集,用于研究两种相互竞争的在线社区:传播虚假信息的误导性用户与反驳虚假信息的知情用户。通过网络分析、社会语言学建模和机器人账号检测,研究发现误导性社区的网络密度更高、组织性更强,且机器人账号比例更高,表明存在有组织的虚假信息传播活动;而知情用户则更多依赖叙事框架来反驳虚假信息。
From conspiracy theories to fake cures and fake treatments, COVID-19 has become a hot-bed for the spread of misinformation online. It is more important than ever to identify methods to debunk and correct false information online. Detection and characterization of misinformation requires an availability of annotated datasets. Most of the published COVID-19 Twitter datasets are generic, lack annotations or labels, employ automated annotations using transfer learning or semi-supervised methods, or are not specifically designed for misinformation. Annotated datasets are either only focused on "fake news", are small in size, or have less diversity in terms of classes. Here, we present a novel Twitter misinformation dataset called <strong>"CMU-MisCov19"</strong> with 4573 annotated tweets over 17 themes around the COVID-19 discourse. We also present our annotation codebook for the different COVID-19 themes on Twitter, along with their descriptions and examples, for the community to use for collecting further annotations. Further details related to the dataset, and our analysis based on this dataset can be found at https://arxiv.org/abs/2008.00791. In adherence to the Twitter’s terms and conditions, we do not provide the full tweet JSONs but provide a ".csv" file with the tweet IDs so that the tweets can be rehydrated. We also provide the annotations, and the date of creation for each tweet for the reproduction of the results of our analyses. <strong>Note: If for any reason, you are not able to rehydrate all the tweets, reach out to Shahan Ali Memon at (shahan@nyu.edu).</strong> If you use this data, please cite our paper as follows: <em>"Shahan Ali Memon and Kathleen M. Carley. Characterizing COVID-19 Misinformation Communities Using a Novel Twitter Dataset, In Proceedings of The 5th International Workshop on Mining Actionable Insights from Social Networks (MAISoN 2020), co-located with CIKM, virtual event due to COVID-19, 2020."</em>
研究动机与目标
- 收集并发布一个多样化、人工标注的Twitter数据集,聚焦于COVID-19虚假信息,供公众研究使用。
- 描述大流行期间误导性与知情在线社区在结构、语言和行为方面的差异。
- 调查机器人账号和社区成员身份在放大虚假信息传播中的作用,特别是在反疫苗网络中。
- 评估基于叙事的传播方式是否比直接反驳更有效地遏制虚假信息。
提出的方法
- 使用基于关键词的搜索API,在三个非连续日期(2020年3月29日、6月15日和6月24日)内收集Twitter数据,重点关注与虚假信息主题相关的术语。
- 使用综合编码手册对推文进行标注,包含多个标注层级,包括虚假信息类型、立场和语言特征。
- 基于转发、提及和回复关系构建用户级网络,以分析社区结构和中心性指标。
- 应用预训练的机器人检测模型(Bot-Hunter),设定置信度阈值≥0.75,以识别自动化账号。
- 使用自然语言处理技术进行社会语言学分析,比较不同社区在情感基调和叙事使用上的差异。
- 映射用户在多个虚假信息子社区中的成员身份,包括反疫苗和其它与健康相关的虚假信息传播网络。
实验结果
研究问题
- RQ1在密度和中心性方面,误导性与知情的COVID-19 Twitter社区的网络结构有何不同?
- RQ2语言模式,特别是叙事使用和情感基调,如何区分知情用户与误导性用户?
- RQ3机器人在虚假信息社区中的占比有多大?其存在是否暗示存在有组织的虚假信息传播活动?
- RQ4误导性用户是否在反疫苗社区中占比较高?
- RQ5基于叙事的传播方式是否比直接反驳更有效地遏制虚假信息?
主要发现
- 误导性社区的网络密度显著高于知情社区,表明其组织更紧密、更 cohesive。
- 误导性社区中机器人比例显著更高(p < 0.05),表明存在有组织的虚假信息传播活动。
- 绝大多数误导性用户属于反疫苗社区,表明反疫苗网络与更广泛的虚假信息传播网络之间存在强烈重叠。
- 知情用户在其帖子中显著使用更多叙事结构,而误导性用户则更多依赖陈述性或阴谋论式言论。
- 两个社区均表现出负面情感基调,但知情用户使用叙事可能反映了一种策略性传播方式。
- 研究结果挑战了‘通用驳斥信息普遍有效’的假设,暗示可能需要采用量身定制的、基于叙事的干预措施。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。