Skip to main content
QUICK REVIEW

[论文解读] Predicting Misinformation and Engagement in COVID-19 Twitter Discourse in the First Months of the Outbreak

Mirela Silva, Fabrício Ceschin|arXiv (Cornell University)|Dec 3, 2020
Misinformation and Its Impacts被引用 13
一句话总结

本研究分析了505,000条与COVID-19相关的推文,利用170多个特征预测虚假信息和传播度,发现文本内容在区分事实与虚假信息方面最为关键,而用户元数据则在预测高传播度方面表现最佳。该研究在十个分类器中均实现了超过72%的F1分数,揭示出机器人账号不成比例地传播虚假信息,而虚假信息的传播度低于真实内容。

ABSTRACT

Disinformation entails the purposeful dissemination of falsehoods towards a greater dubious agenda and the chaotic fracturing of a society. The general public has grown aware of the misuse of social media towards these nefarious ends, where even global public health crises have not been immune to misinformation (deceptive content spread without intended malice). In this paper, we examine nearly 505K COVID-19-related tweets from the initial months of the pandemic to understand misinformation as a function of bot-behavior and engagement. Using a correlation-based feature selection method, we selected the 11 most relevant feature subsets among over 170 features to distinguish misinformation from facts, and to predict highly engaging misinformation tweets about COVID-19. We achieved an average F-score of at least 72\% with ten popular multi-class classifiers, reinforcing the relevance of the selected features. We found that (i) real users tweet both facts and misinformation, while bots tweet proportionally more misinformation; (ii) misinformation tweets were less engaging than facts; (iii) the textual content of a tweet was the most important to distinguish fact from misinformation while (iv) user account metadata and human-like activity were most important to predict high engagement in factual and misinformation tweets; and (v) sentiment features were not relevant.

研究动机与目标

  • 理解在早期COVID-19大流行期间,机器人行为与用户传播度在虚假信息传播中的作用。
  • 识别在Twitter话语中,哪些特征最能有效区分虚假信息与真实内容。
  • 预测哪些推文(无论是真实还是误导性)会获得高传播度。
  • 评估情感特征在区分虚假信息和预测病毒式传播中的相关性。

提出的方法

  • 本研究分析了大流行初期数月内的近505,000条与COVID-19相关的推文。
  • 采用基于相关性的特征选择方法,从170多个候选特征中识别出11个最相关的特征子集。
  • 使用选定特征训练并评估了十个多分类器,以预测虚假信息和高传播度。
  • 对文本内容、用户账户元数据和行为模式(例如类人活动)进行分析,作为关键预测因子。
  • F1分数被用作主要评估指标,以衡量各类分类器的性能表现。
  • 情感特征被明确测试,结果发现其在两项分类任务中均无相关性。

实验结果

研究问题

  • RQ1在早期COVID-19的Twitter话语中,哪些特征最能有效区分虚假信息与真实内容?
  • RQ2机器人行为在传播虚假信息与真实信息方面与人类用户有何不同?
  • RQ3哪些因素能预测真实与虚假信息推文的高传播度?
  • RQ4情感特征在预测虚假信息或传播度方面有多大的贡献?

主要发现

  • 真实用户既传播事实也传播虚假信息,而机器人账号则不成比例地传播虚假信息。
  • 虚假信息推文的传播度低于真实内容推文,表明尽管机器人传播更广泛,但其病毒式传播性较低。
  • 文本内容是区分虚假信息与事实的最重要特征,其表现优于元数据和行为信号。
  • 用户账户元数据和类人活动模式是预测高传播度推文的最具预测力的特征。
  • 情感特征在虚假信息分类或传播度预测中均无显著相关性。
  • 所选特征子集使十个多分类器在两项预测任务中均实现了至少72%的平均F1分数。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。