[论文解读] False News On Social Media: A Data-Driven Survey
这项以数据为驱动的综述综合分析了近年来通过分析跨多样化数据集的文本、行为和网络特征,在社交媒体上检测、表征和缓解虚假新闻的进展。它将网络动态、心理偏见和可扩展的检测技术确定为应对错误信息的核心,其关键贡献在于建模谣言传播机制,并利用众包事实核查实现现实世界的影响。
In the past few years, the research community has dedicated growing interest to the issue of false news circulating on social networks. The widespread attention on detecting and characterizing false news has been motivated by considerable backlashes of this threat against the real world. As a matter of fact, social media platforms exhibit peculiar characteristics, with respect to traditional news outlets, which have been particularly favorable to the proliferation of deceptive information. They also present unique challenges for all kind of potential interventions on the subject. As this issue becomes of global concern, it is also gaining more attention in academia. The aim of this survey is to offer a comprehensive study on the recent advances in terms of detection, characterization and mitigation of false news that propagate on social media, as well as the challenges and the open questions that await future research on the field. We use a data-driven approach, focusing on a classification of the features that are used in each study to characterize false information and on the datasets used for instructing classification methods. At the end of the survey, we highlight emerging approaches that look most promising for addressing false news.
研究动机与目标
- 提供对社交媒体平台上虚假新闻检测、表征与缓解领域近期研究的全面、数据驱动的综述。
- 识别并分类虚假新闻检测研究中使用的关键特征——文本、行为和基于网络的特征。
- 评估各研究中采用的数据集和评估方法论,以评估可复现性和可比性。
- 突出有前景的缓解策略,特别是利用众包事实核查和用户主动参与的策略。
- 概述开放挑战与未来研究方向,包括跨学科协作以及检测系统的现实世界部署。
提出的方法
- 系统性地调研了190余篇关于社交媒体中虚假新闻的研究,重点关注检测、表征与缓解技术。
- 将检测中使用的特征分类为文本特征(例如语言线索)、行为特征(例如分享模式)和基于网络的特征(例如传播结构)三类。
- 分析Rumors、Hoaxy和基于Twitter的数据集,以评估检测模型的性能和泛化能力。
- 使用精确率、错误信息减少量和暴露预防等指标评估检测方法,特别是在时间序列和随机建模框架中。
- 研究主动缓解策略,包括算法选择需核查内容(例如CURB和自适应标记模型)。
- 整合心理学(例如确认偏误、天真现实主义)和社会网络科学的见解,以指导检测与干预设计。
实验结果
研究问题
- RQ1在社交媒体中,区分虚假新闻与真实新闻的主导文本、行为和基于网络的特征是什么?
- RQ2不同的数据集和评估协议在多大程度上影响虚假新闻检测模型的性能和可比性?
- RQ3众包事实核查和用户标记在多大程度上能有效减少错误信息的传播?
- RQ4社交机器人、回音室和网络结构在放大虚假新闻传播中扮演什么角色?
- RQ5构建可扩展、现实世界部署的缓解系统的关键开放挑战与未来研究方向是什么?
主要发现
- 虚假新闻在社交媒体上传播速度更快、范围更广,网络结构和社交机器人在放大过程中起着关键作用。
- 文本特征如情感语言和语言复杂性虽具参考价值,但单独使用效果有限;社会和网络特征可显著提升检测准确率。
- 众包事实核查机制(如CURB模型和自适应标记算法)可在理论保证下减少错误信息暴露。
- 考虑用户可信度和对抗性行为(如垃圾信息用户)的模型在模拟环境中展现出更高的鲁棒性和性能。
- 缺乏标准化的黄金标准数据集和评估协议,阻碍了该领域研究的可复现性和跨研究比较。
- 未来的检测系统必须整合心理学、新闻学和计算机科学的洞见,以构建更有效、可扩展且可现实部署的解决方案。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。