Skip to main content
QUICK REVIEW

[论文解读] Facebook Ad Engagement in the Russian Active Measures Campaign of 2016

Mirela Silva, Luiz Giovanini|arXiv (Cornell University)|Dec 21, 2020
Misinformation and Its Impacts参考文献 30被引用 4
一句话总结

本研究利用机器学习方法,分析了2016年美国大选期间俄罗斯互联网研究机构(IRA)发布的3,517则Facebook广告,识别出影响用户参与度的关键特征。研究发现,广告支出、文本大小、广告持续时间以及积极情绪是参与度最强的预测因素,且与社会语言学特征(如与宗教相关的语言)也密切相关,这些特征在有效传播虚假信息的活动中起到关键作用。

ABSTRACT

This paper examines 3,517 Facebook ads created by Russia's Internet Research Agency (IRA) between June 2015 and August 2017 in its active measures disinformation campaign targeting the 2016 U.S. general election. We aimed to unearth the relationship between ad engagement (as measured by ad clicks) and 41 features related to ads' metadata, sociolinguistic structures, and sentiment. Our analysis was three-fold: (i) understand the relationship between engagement and features via correlation analysis; (ii) find the most relevant feature subsets to predict engagement via feature selection; and (iii) find the semantic topics that best characterize the dataset via topic modeling. We found that ad expenditure, text size, ad lifetime, and sentiment were the top features predicting users' engagement to the ads. Additionally, positive sentiment ads were more engaging than negative ads, and sociolinguistic features (e.g., use of religion-relevant words) were identified as highly important in the makeup of an engaging ad. Linear SVM and Logistic Regression classifiers achieved the highest mean F-scores (93.6% for both models), determining that the optimal feature subset contains 12 and 6 features, respectively. Finally, we corroborate the findings of related works that the IRA specifically targeted Americans on divisive ad topics (e.g., LGBT rights, African American reparations).

研究动机与目标

  • 识别2016年美国大选期间俄罗斯IRA Facebook广告中影响用户参与度的特征。
  • 理解社会语言学特征、情感和元数据在塑造广告有效性方面的作用。
  • 评估机器学习模型在识别高参与度虚假信息内容方面的预测能力。
  • 通过广告内容的主题建模,揭示IRA所针对的社会政治议题。
  • 通过揭示国家支持的虚假信息活动中的结构与语言模式,为未来反制措施提供依据。

提出的方法

  • 从3,517则IRA Facebook广告中提取了41项特征,包括元数据(如广告支出、持续时间)、文本大小和情感分析结果。
  • 通过相关性分析评估参与度(点击量)与各项特征之间的关系。
  • 在六种机器学习模型(如线性SVM、逻辑回归)中应用特征选择技术,以识别最优特征子集。
  • 采用潜在狄利克雷分布(LDA)进行主题建模,以识别广告内容中的主要语义主题。
  • 使用词袋(BoW)表示法作为LDA的文本输入,并采用17种LIWC类别进行社会语言学分析。
  • 训练并评估多种分类器以确定其预测性能,其中线性SVM和逻辑回归模型的F-score最高。

实验结果

研究问题

  • RQ1哪些元数据、情感和社交语言学特征与IRA Facebook广告的参与度最强相关?
  • RQ2哪些特征子集能最优预测虚假信息广告的高参与度?
  • RQ3IRA在其虚假信息活动中最常针对哪些社会政治议题?
  • RQ4基于提取的特征,机器学习模型在分类高参与度虚假信息广告方面的表现如何?
  • RQ5情感和语言模式(如与宗教相关的词汇)在多大程度上影响广告的有效性?

主要发现

  • 广告支出、文本大小、广告持续时间以及情感是六种机器学习模型一致识别出的、对参与度最具预测力的前四项特征。
  • 积极情绪广告的参与度显著高于消极情绪广告,表明情感基调是关键的参与度驱动因素。
  • 社会语言学特征——尤其是与宗教相关的词汇使用——是参与度最重要的预测因素之一。
  • 线性SVM和逻辑回归模型的平均F-score最高,达到93.6%,其最优特征子集分别包含12个和6个特征。
  • 主题建模结果显示,IRA特别针对诸如LGBT权利、非裔美国人赔偿、第二修正案权利以及生殖权利等具有分裂性的社会政治议题。
  • 本研究证实了既往发现:IRA在2016年大选周期中战略性地针对边缘化和极化的社群,以加剧社会分裂。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。