Skip to main content
QUICK REVIEW

[论文解读] Intersectional Bias in Hate Speech and Abusive Language Datasets

Jae‐Yeon Kim, Carlos Ortiz|arXiv (Cornell University)|May 12, 2020
Hate Speech and Cyberbullying Detection参考文献 24被引用 16
一句话总结

本研究通过分析来自公开Twitter数据集的99,996条推文中的种族、性别和政治归属,调查了仇恨言论和辱骂性语言数据集中存在的交叉性偏见。研究发现,非裔美国人推文——尤其是非裔美国黑人男性——被标记为辱骂或仇恨言论的可能性显著更高,其被标记为辱骂的几率最高达3.7倍,被标记为仇恨言论的几率高出77%,即使在控制了政党认同后依然如此,这是首次系统性地揭示此类数据集中存在的交叉性偏见。

ABSTRACT

Algorithms are widely applied to detect hate speech and abusive language in social media. We investigated whether the human-annotated data used to train these algorithms are biased. We utilized a publicly available annotated Twitter dataset (Founta et al. 2018) and classified the racial, gender, and party identification dimensions of 99,996 tweets. The results showed that African American tweets were up to 3.7 times more likely to be labeled as abusive, and African American male tweets were up to 77% more likely to be labeled as hateful compared to the others. These patterns were statistically significant and robust even when party identification was added as a control variable. This study provides the first systematic evidence on intersectional bias in datasets of hate speech and abusive language.

研究动机与目标

  • 调查仇恨言论和辱骂性语言数据集是否在种族、性别和政治归属等交叉社会身份上表现出偏见。
  • 确定与其它人口群体相比,非裔美国人推文是否被不成比例地标记为辱骂或仇恨言论。
  • 评估此类偏见在控制政党认同后是否依然存在。
  • 为广泛使用的自然语言处理(NLP)仇恨言论检测数据集中存在的交叉性偏见提供系统性实证证据。

提出的方法

  • 本研究分析了一个公开可用的Twitter数据集(Founta et al., 2018),其中包含99,996条人工标注的推文。
  • 通过语言线索和上下文分析相结合的方法,对每条推文的种族、性别和政党身份进行分类。
  • 应用统计建模方法,检验人口身份与标注结果(辱骂、仇恨)之间的关联。
  • 使用多元回归模型控制政党认同,以隔离种族和性别对标注结果的影响。
  • 包括稳健性检验,以确保研究结果并非模型设定或数据不平衡的产物。
  • 在交叉性类别中评估结果的统计显著性与效应量。

实验结果

研究问题

  • RQ1在仇恨言论数据集中,与其它种族群体相比,非裔美国人推文是否更可能被标记为辱骂或仇恨言论?
  • RQ2作为黑人男性,其身份是否会使被标记为仇恨言论的可能性相较于其它人口群体显著提高?
  • RQ3在控制政党认同后,观察到的标注偏差是否依然稳健?
  • RQ4交叉性身份(种族、性别、政党)在多大程度上导致了标注结果的差异?

主要发现

  • 非裔美国人推文被标记为辱骂的可能性最高达其它种族群体的3.7倍。
  • 非裔美国黑人男性推文被标记为仇恨言论的可能性比其它人口群体高出最多77%。
  • 即使在控制政党认同后,这些差异依然具有统计显著性。
  • 该偏见在多个交叉性身份组合中保持一致,表明标注结果存在系统性偏差。
  • 本研究首次系统性地提供了仇恨言论和辱骂性语言数据集中存在交叉性偏见的证据。
  • 研究结果表明,用于训练NLP模型的数据可能加剧并放大自动化内容审核中的社会不公。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。