Skip to main content
QUICK REVIEW

[论文解读] The COVID-19 Infodemic: Can the Crowd Judge Recent Misinformation Objectively?

Kevin Roitero, Michael Soprano|arXiv (Cornell University)|Aug 13, 2020
Misinformation and Its Impacts参考文献 33被引用 4
一句话总结

本研究调查了非专家众包工作者是否能客观评估近期与新冠相关的虚假信息的真实性。通过定制化的搜索界面和行为日志记录,作者发现众包可产生可靠的真伪判断,尤其当工作者从选定网络来源复制文本时;此外,工作者的背景和行为与判断质量显著相关,为改进聚合方法提供了新信号。

ABSTRACT

Misinformation is an ever increasing problem that is difficult to solve for the research community and has a negative impact on the society at large. Very recently, the problem has been addressed with a crowdsourcing-based approach to scale up labeling efforts: to assess the truthfulness of a statement, instead of relying on a few experts, a crowd of (non-expert) judges is exploited. We follow the same approach to study whether crowdsourcing is an effective and reliable method to assess statements truthfulness during a pandemic. We specifically target statements related to the COVID-19 health emergency, that is still ongoing at the time of the study and has arguably caused an increase of the amount of misinformation that is spreading online (a phenomenon for which the term "infodemic" has been used). By doing so, we are able to address (mis)information that is both related to a sensitive and personal issue like health and very recent as compared to when the judgment is done: two issues that have not been analyzed in related work. In our experiment, crowd workers are asked to assess the truthfulness of statements, as well as to provide evidence for the assessments as a URL and a text justification. Besides showing that the crowd is able to accurately judge the truthfulness of the statements, we also report results on many different aspects, including: agreement among workers, the effect of different aggregation functions, of scales transformations, and of workers background / bias. We also analyze workers behavior, in terms of queries submitted, URLs found / selected, text justifications, and other behavioral data like clicks and mouse actions collected by means of an ad hoc logger.

研究动机与目标

  • 评估非专家众包工作者是否能客观判断在持续的新冠疫情期间,近期敏感的健康相关虚假信息的真实性。
  • 分析工作者背景、认知行为及信息检索模式如何影响真伪判断的质量。
  • 探究行为信号(如URL选择、文本解释模式和交互日志)是否能预测工作者准确性并改进聚合方法。
  • 评估众包标签相对于专家标签的可靠性,特别是针对真伪量表两端的极端类别(如“完全错误”和“虚假”)。
  • 发布收集到的数据集和行为日志,以支持未来虚假信息检测与真伪评估的研究。

提出的方法

  • 众包工作者被要求使用定制开发的搜索界面,对8条近期与新冠相关的陈述进行真伪判断,系统记录了所有交互行为,包括点击、鼠标移动和输入行为。
  • 工作者在5分制上提供真伪评分,并附上URL和文本解释,且必须引用网络来源的证据。
  • 本研究采用行为日志技术,捕获详细的交互数据,包括解释中是否从选定URL复制了文本。
  • 通过将工作者的判断与经专家验证的标签进行比较,评估工作者质量,准确性以预测误差和CEM${}^{ ext{ORD}}$得分衡量。
  • 测试了聚合函数与量表变换以提升标签质量,并利用工作者的政治自我认同评估偏见影响。
  • 统计分析关联了行为模式(如从来源复制文本、自由文本使用)与判断准确性及误差分布。

实验结果

研究问题

  • RQ1非专家众包工作者能否客观评估新冠疫情期间近期健康相关虚假信息的真实性?
  • RQ2工作者背景与自我报告的政治取向如何与判断质量相关?
  • RQ3哪些行为信号(如URL选择、文本复制或解释模式)能预测工作者准确性?
  • RQ4众包判断在多大程度上能与经专家验证的真伪标签一致,特别是针对“完全错误”或“虚假”等极端类别?
  • RQ5点击、鼠标操作和文本解释模式等行为数据能否用于改进未来真伪标注的聚合方法?

主要发现

  • 众包工作者在客观评估近期新冠相关陈述真伪方面表现出色,聚合众包标签与专家标签高度一致,仅在两个最极端类别(“完全错误”和“虚假”)存在例外。
  • 在解释中从选定URL复制文本的工作者,其预测误差显著更低,CEM${}^{ ext{ORD}}$得分更高(0.62 vs. 0.58),表明判断质量更优。
  • 预测误差分布呈非对称性,对真伪的高估更为常见(正误差),尤其在被标记为“虚假”或“完全错误”的陈述中。
  • 工作者使用了多种来源,包括事实核查网站和健康相关网站,其解释行为与判断准确性高度相关。
  • 工作者自我报告的政治背景被发现与标签质量相关,提示人口统计因素会影响判断可靠性。
  • 文本复制和来源引用模式等行为信号,为改进众包真伪评估中的未来聚合方法提供了有前景且可量化的指标。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。