Skip to main content
QUICK REVIEW

[Paper Review] Hate Speech Classifiers Learn Human-Like Social Stereotypes

Aida Mostafazadeh Davani, Mohammad Atari|arXiv (Cornell University)|Oct 28, 2021
Hate Speech and Cyberbullying Detection37 references4 citations
TL;DR

This study demonstrates that hate speech classifiers learn and replicate human-like social stereotypes, leading to biased predictions that disproportionately target marginalized groups. By combining social psychology (IAT) and NLP, the authors show that both annotator implicit biases and language-embedded stereotypes correlate with systematic misclassification in neural hate speech detectors, revealing a critical source of unfairness in AI moderation systems.

ABSTRACT

Social stereotypes negatively impact individuals' judgements about different groups and may have a critical role in how people understand language directed toward minority social groups. Here, we assess the role of social stereotypes in the automated detection of hateful language by examining the relation between individual annotator biases and erroneous classification of texts by hate speech classifiers. Specifically, in Study 1 we investigate the impact of novice annotators' stereotypes on their hate-speech-annotation behavior. In Study 2 we examine the effect of language-embedded stereotypes on expert annotators' aggregated judgements in a large annotated corpus. Finally, in Study 3 we demonstrate how language-embedded stereotypes are associated with systematic prediction errors in a neural-network hate speech classifier. Our results demonstrate that hate speech classifiers learn human-like biases which can further perpetuate social inequalities when propagated at scale. This framework, combining social psychological and computational linguistic methods, provides insights into additional sources of bias in hate speech moderation, informing ongoing debates regarding fairness in machine learning.

Motivation & Objective

  • To investigate how individual annotators' implicit social stereotypes influence their hate speech labeling behavior.
  • To examine whether language-embedded stereotypes in text correlate with higher inter-annotator disagreement among expert annotators.
  • To assess the impact of social stereotypes on systematic prediction errors in neural hate speech classifiers.
  • To identify how human-like biases in data and annotation processes are internalized by AI models, exacerbating social inequities.

Proposed method

  • Administered an Implicit Association Test (IAT) to quantify participants' implicit biases toward eight social group pairs (e.g., Black vs White, Gay vs Straight).
  • Collected hate speech annotations from 124 annotators on a diverse corpus of social media posts, including those mentioning specific social groups.
  • Applied a multi-level Poisson regression model to analyze the relationship between IAT scores (implicit bias) and the number of hate speech labels assigned per social group.
  • Conducted two-sample permutation tests to compare inter-annotator disagreement on posts with and without social group tokens, and on hateful vs non-hateful content.
  • Performed a two-way ANOVA to examine the interaction effect of hate speech label and social group token presence on disagreement scores.
  • Trained a neural network hate speech classifier and analyzed its prediction errors in relation to stereotypical language patterns and group mentions.

Experimental results

Research questions

  • RQ1How do annotators' implicit biases, measured via IAT, correlate with the number of hate speech labels they assign to content mentioning specific social groups?
  • RQ2To what extent do mentions of social groups in text increase inter-annotator disagreement among expert annotators?
  • RQ3Does the presence of both hate speech content and social group tokens in a post lead to higher or lower disagreement among annotators?
  • RQ4To what extent do language-embedded social stereotypes contribute to systematic errors in neural hate speech classifiers?

Key findings

  • A one-point increase in an annotator’s implicit bias score for a social group is associated with a 0.1% increase in the number of hate speech labels assigned to content mentioning that group (β = 0.09, p < .01).
  • Inter-annotator disagreement is significantly higher on posts labeled as hate speech (M = 0.50) than on non-hateful posts (M = 0.13), with p < .001.
  • Posts mentioning social group tokens trigger significantly higher disagreement (M = 0.30) compared to those without (M = 0.13), with p < .001.
  • The two-way ANOVA reveals that both the presence of hate speech and social group tokens independently predict higher disagreement (p < .001 for both).
  • The neural hate speech classifier exhibits systematic prediction errors that align with stereotypical associations in language, indicating that models learn and amplify human-like social biases.
  • The study confirms that both annotator-level implicit biases and language-embedded stereotypes contribute to unfair model behavior, reinforcing social inequalities at scale.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.