Skip to main content
QUICK REVIEW

[论文解读] A curated collection of COVID-19 online datasets

Isa Inuwa-Dutse, Ioannis Korkontzelos|arXiv (Cornell University)|Jul 19, 2020
Misinformation and Its Impacts参考文献 12被引用 6
一句话总结

本文呈现了一个精心整理的1,169个与COVID-19相关的数据集,涵盖Twitter、可信健康来源及世卫组织全球疫情报告,支持对信息过载现象的缓解研究。这些数据集支持内容分类、真实性验证、语气分析、错误信息传播、主题分析以及政策情绪追踪,提供完整的数据还原支持,并为事实核查系统提供基准测试。

ABSTRACT

One of the defining moments of the year 2020 is the outbreak of Coronavirus Disease (Covid-19), a deadly virus affecting the body's respiratory system to the point of needing a breathing aid via ventilators. As of June 21, 2020 there are 12,929,306 confirmed cases and 569,738 confirmed deaths across 216 countries, areas or territories. The scale of spread and impact of the pandemic left many nations grappling with preventive and curative approaches. The infamous lockdown measure introduced to mitigate the virus spread has altered many aspects of our social routines in which demand for online-based services skyrocketed. As the virus propagate, so does misinformation and fake news around it via online social media, which seems to favour virality over veracity. With a majority of the populace confined to their homes for a long period, vulnerability to the toxic impact of online misinformation is high. A case in point is the various myths and disinformation associated with the Covid-19, which, if left unchecked, could lead to a catastrophic outcome and hamper the fight against the virus. While the scientific community is actively engaged in identifying the virus treatment, there is a growing interest in combating the associated harmful infodemic. To this end, researchers have been curating and documenting various datasets about Covid-19. In line with existing studies, we provide an expansive collection of curated datasets to support the fight against the pandemic, especially concerning misinformation. The collection consists of 3 categories of Twitter data, information about standard practices from credible sources and a chronicle of global situation reports. We describe how to retrieve the hydrated version of the data and proffer some research problems that could be addressed using the data.

研究动机与目标

  • 为解决在COVID-19疫情期间研究错误信息与虚假新闻时缺乏真实标签、精心整理的数据集的问题。
  • 通过提供来自多样化来源的结构化、可访问且已还原的结构化数据,支持对信息过载动态的计算研究。
  • 支持开发和基准测试用于检测、分类和验证社交媒体中错误信息的系统。
  • 促进对与疫情相关话题的公众话语进行纵向分析和主题分析,包括政策情绪与社区结构。
  • 支持在危机背景下创建可靠的事实核查与错误信息抑制基准。

提出的方法

  • 收集三类主要数据:基于标签、账号和时间过滤的Twitter数据集、来自可信来源的经核实的健康指南,以及世卫组织全球疫情报告。
  • 使用Twitter官方API检索并还原原始推文ID,恢复完整推文内容以供分析。
  • 应用数据整理技术,确保数据集之间的一致性、相关性及真实标签对齐。
  • 根据内容倾向性将用户分类为支持世卫组织和反对世卫组织的群体,以支持社区检测与情绪分析。
  • 整合外部事实核查资源(例如IFCN、AFP),以验证推文真实性并支持基准测试。
  • 使用流式API收集实时数据,随后进行过滤与主题聚类,聚焦于相关疫情话题。

实验结果

研究问题

  • RQ1如何训练内容分类模型,以区分社交媒体上真实的疫情信息与错误信息?
  • RQ2利用权威健康来源的真实标签数据,验证推文真实性的最有效方法是什么?
  • RQ3错误信息的语气(如情绪化、负面或耸动)与事实性内容有何不同?
  • RQ4在疫情期间,错误信息与准确信息的传播在结构和行为模式上存在哪些差异?
  • RQ5公众对封锁政策的情绪随时间如何演变?情绪是否可被可靠地与政策措施关联?

主要发现

  • 该数据集集合包含来自Twitter、世卫组织报告及经核实的健康来源的1,169个精心整理的数据集,为信息过载研究提供了全面基础。
  • 包含已还原的推文数据,支持全文分析,克服了仅含ID的原始数据集的局限性。
  • 该数据集支持识别出不同的用户群体,包括支持和反对世卫组织的群体,且可能存在重叠的子群体。
  • 数据支持纵向情绪分析,显示公众对疫情措施的感知随时间的变化。
  • 整合权威来源使得事实核查与错误信息检测系统的基准测试成为可能。
  • 该数据集通过支持围绕特定主题的推文过滤与聚类,促进主题分析,提升上下文理解能力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。