Skip to main content
QUICK REVIEW

[论文解读] Towards Measuring Adversarial Twitter Interactions against Candidates in the US Midterm Elections

Yiqing Hua, Thomas Ristenpart|arXiv (Cornell University)|May 9, 2020
Hate Speech and Cyberbullying Detection参考文献 29被引用 13
一句话总结

本文提出一种上下文感知方法,用于检测针对2018年中期选举期间美国众议院候选人的对抗性互动——如有毒回复、身份攻击和消息重复。通过结合Perspective API、政党倾向启发式方法以及一种新颖的标签传播算法以构建候选人的专属对抗性词表,该方法提升了检测精度,并揭示了通用毒性模型难以捕捉的细微、上下文相关的滥用行为,分析了来自786名候选人的170万条推文。

ABSTRACT

Adversarial interactions against politicians on social media such as Twitter have significant impact on society. In particular they disrupt substantive political discussions online, and may discourage people from seeking public office. In this study, we measure the adversarial interactions against candidates for the US House of Representatives during the run-up to the 2018 US general election. We gather a new dataset consisting of 1.7 million tweets involving candidates, one of the largest corpora focusing on political discourse. We then develop a new technique for detecting tweets with toxic content that are directed at any specific candidate.Such technique allows us to more accurately quantify adversarial interactions towards political candidates. Further, we introduce an algorithm to induce candidate-specific adversarial terms to capture more nuanced adversarial interactions that previous techniques may not consider toxic. Finally, we use these techniques to outline the breadth of adversarial interactions seen in the election, including offensive name-calling, threats of violence, posting discrediting information, attacks on identity, and adversarial message repetition.

研究动机与目标

  • 解决现有上下文无关毒性检测在识别针对政治候选人推特上的对抗性互动时的局限性。
  • 通过整合政党倾向与互动上下文,提升定向毒性检测的精度,以推断有毒内容的真实目标。
  • 发现捕捉细微、个性化形式滥用的候选人专属对抗性词表,这些形式未被通用语言模型识别。
  • 量化对抗性互动的广度与类型,包括身份攻击、威胁、虚假信息和消息重复。
  • 提供一种可扩展、高精度的框架,用于检测政治话语中的对抗性内容,支持更健康的在线民主参与。

提出的方法

  • 结合Perspective API(一种先进的基于语言的毒性分类器)与启发式方法,推断用户的政治倾向,从而确定有毒回复的可能目标。
  • 利用互动上下文(如提及、回复和转发)判断有毒内容是否针对特定候选人,减少非定向滥用带来的误报。
  • 在政治网络图上应用标签传播算法,发现候选人专属的对抗性词表,识别在特定语境中具有攻击性但本身并非明显有毒的短语。
  • 利用美国政治的二元结构(共和党对民主党)作为代理,提升在模糊互动中目标推断的准确性。
  • 采用多阶段标注流程,由人工标注员验证检测结果并分类对抗性互动类型。
  • 分析包含786名候选人、共170万条推文的数据集,重点关注毒性、身份攻击、威胁、虚假信息及消息重复模式。

实验结果

研究问题

  • RQ1当毒性内容并不总是针对提及或回复的用户时,如何更准确地检测推特上的对抗性互动?
  • RQ2政党倾向在提升政治候选人定向毒性检测精度方面发挥何种作用?
  • RQ3哪些类型的对抗性互动被上下文无关的语言模型所遗漏,但可通过候选人专属词表被发现?
  • RQ4候选人属性(如性别与政党归属)与所接收对抗性互动数量之间存在何种关联?
  • RQ5对抗性消息重复与基于身份的攻击在多大程度上影响了政治推特话语的整体毒性格局?

主要发现

  • 候选人从推特获得的关注总量是预测对抗性互动的最强指标,性别与政党归属无显著影响。
  • 结合政党倾向启发式方法与Perspective API可显著提升检测精度,减少非定向滥用带来的误报,尤其在回复与提及场景中。
  • 所提出的标签传播算法成功识别出候选人专属的对抗性词表——如个性化侮辱或编码化蔑称——这些短语未被通用模型标记为有毒。
  • 大量对抗性互动涉及基于身份的攻击,包括厌女言论和种族化蔑称,需依赖上下文理解才能检测。
  • 消息重复(如反复呼吁辞职或编造丑闻)是一种常见对抗性策略,若不分析推文序列则难以检测。
  • 人工标注员发现,许多对抗性互动需要背景知识与上下文解读,凸显即使使用先进模型,自动化检测仍存在局限。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。