Skip to main content
QUICK REVIEW

[论文解读] Detecting Propagators of Disinformation on Twitter Using Quantitative Discursive Analysis

Mark M. Bailey|arXiv (Cornell University)|Oct 11, 2022
Misinformation and Its Impacts被引用 4
一句话总结

本文提出一种定量话语分析方法,利用中心化共振分析和Clauset-Newman-Moore社区检测,在2016年美国大选期间检测Twitter上的俄罗斯虚假信息机器人。该方法在分类机器人时实现了0.9070的Matthews相关系数(MCC),表现出高敏感性,表明传播者之间存在强烈的话语相似性,但因对照组中话语不相似,难以识别非机器人。

ABSTRACT

Efforts by foreign actors to influence public opinion have gained considerable attention because of their potential to impact democratic elections. Thus, the ability to identify and counter sources of disinformation is increasingly becoming a top priority for government entities in order to protect the integrity of democratic processes. This study presents a method of identifying Russian disinformation bots on Twitter using centering resonance analysis and Clauset-Newman-Moore community detection. The data reflect a significant degree of discursive dissimilarity between known Russian disinformation bots and a control set of Twitter users during the timeframe of the 2016 U.S. Presidential Election. The data also demonstrate statistically significant classification capabilities (MCC = 0.9070) based on community clustering. The prediction algorithm is very effective at identifying true positives (bots), but is not able to resolve true negatives (non-bots) because of the lack of discursive similarity between control users. This leads to a highly sensitive means of identifying propagators of disinformation with a high degree of discursive similarity on Twitter, with implications for limiting the spread of disinformation that could impact democratic processes.

研究动机与目标

  • 识别2016年美国总统大选期间Twitter上虚假信息的传播者。
  • 解决利用话语模式区分虚假信息机器人与合法用户的问题。
  • 开发一种结合语言与网络结构的方法,以检测有组织的虚假信息传播活动。
  • 评估定量话语分析在高敏感性分类机器人行为方面的有效性。

提出的方法

  • 应用中心化共振分析以建模用户生成内容中的话语连贯性与焦点。
  • 采用Clauset-Newman-Moore社区检测识别具有相似话语模式的用户群集。
  • 量化已知俄罗斯虚假信息机器人与非机器人用户对照组之间的话语相似性。
  • 通过比较用户间的话语轨迹,检测表现出高度相似性的社区,以识别有组织的机器人活动。
  • 基于社区群集训练分类模型,依据话语结构预测机器人身份。
  • 该方法依赖于语言连贯性与网络聚类的统计分析,以识别虚假信息传播者。

实验结果

研究问题

  • RQ1话语相似性模式能否有效区分2016年美国大选期间Twitter上俄罗斯虚假信息机器人与非机器人用户?
  • RQ2已知的虚假信息机器人与用户对照组相比,其话语连贯性在统计上是否显著?
  • RQ3基于话语模式的社区聚类在分类机器人账户方面有多有效?
  • RQ4为何该模型对真正阳性(即实际机器人)具有高敏感性,却无法在对照组中识别真正阴性?
  • RQ5这些发现对检测和缓解有组织的虚假信息传播活动具有何种启示?

主要发现

  • 该方法实现了0.9070的Matthews相关系数(MCC),表明分类性能优异。
  • 在已知俄罗斯虚假信息机器人与非机器人用户对照组之间观察到显著的话语不相似性。
  • 由于机器人群集中话语相似性强烈,模型在识别真正阳性(即实际机器人)方面表现出高敏感性。
  • 无法识别真正阴性的原因在于对照组用户之间缺乏话语相似性,从而限制了负样本分类能力。
  • 结果表明,有组织机器人网络中的话语连贯性可实现对虚假信息传播者的高效检测。
  • 本研究证实,定量话语分析可作为识别虚假信息传播者的关键工具,尤其在高相似性群集中可实现极低的误报率。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。