Skip to main content
QUICK REVIEW

[Paper Review] Is Your Toxicity My Toxicity? Exploring the Impact of Rater Identity on Toxicity Annotation

Nitesh Goyal, Ian Kivlichan|arXiv (Cornell University)|May 1, 2022
Hate Speech and Cyberbullying Detection5 citations
TL;DR

This study investigates how rater identity affects toxicity annotation in online comments, finding that specialized rater pools—composed of individuals identifying as African American or LGBTQ—produce significantly different toxicity ratings compared to a control pool. Models trained on annotations from identity-specific raters outperform those trained on random raters, even with smaller datasets, suggesting that community-identified annotators yield more nuanced, inclusive, and less biased model outcomes.

ABSTRACT

Machine learning models are commonly used to detect toxicity in online conversations. These models are trained on datasets annotated by human raters. We explore how raters' self-described identities impact how they annotate toxicity in online comments. We first define the concept of specialized rater pools: rater pools formed based on raters' self-described identities, rather than at random. We formed three such rater pools for this study--specialized rater pools of raters from the U.S. who identify as African American, LGBTQ, and those who identify as neither. Each of these rater pools annotated the same set of comments, which contains many references to these identity groups. We found that rater identity is a statistically significant factor in how raters will annotate toxicity for identity-related annotations. Using preliminary content analysis, we examined the comments with the most disagreement between rater pools and found nuanced differences in the toxicity annotations. Next, we trained models on the annotations from each of the different rater pools, and compared the scores of these models on comments from several test sets. Finally, we discuss how using raters that self-identify with the subjects of comments can create more inclusive machine learning models, and provide more nuanced ratings than those by random raters.

Motivation & Objective

  • To examine whether rater self-identified identity influences toxicity annotations in online comments.
  • To assess whether specialized rater pools—based on identity groups like African American and LGBTQ—produce more nuanced and less biased annotations than random raters.
  • To evaluate whether models trained on annotations from identity-specific raters outperform those trained on random annotator data.
  • To create and release a publicly available dataset of 382,500 annotations from three curated rater pools for future research.
  • To advocate for integrating identity-based annotator pools into machine learning pipelines to reduce bias in toxicity detection.

Proposed method

  • Three specialized rater pools were formed: self-identified African American, LGBTQ, and neither—each annotating the same 25,500 comments from the Civil Comments dataset.
  • Annotations were collected using a controlled, identity-based rater pool framework, ensuring demographic self-identification as the primary grouping criterion.
  • Statistical analysis was used to compare toxicity ratings across rater pools, identifying significant differences based on rater identity.
  • Content analysis was performed on comments with high disagreement between pools to identify linguistic and contextual markers of disagreement.
  • Multiple machine learning models were trained on annotations from each rater pool and evaluated on various test sets to compare performance.
  • The study released a new dataset of 382,500 annotations from the three specialized rater pools to enable replication and further research.

Experimental results

Research questions

  • RQ1Does the self-identified identity of raters significantly affect their toxicity annotations for comments related to race and sexuality?
  • RQ2How do toxicity ratings from identity-specific rater pools (African American, LGBTQ) differ from those of a control pool (neither African American nor LGBTQ)?
  • RQ3Can models trained on annotations from specialized rater pools outperform models trained on data from random raters, even with smaller datasets?
  • RQ4What linguistic or contextual features explain the differences in ratings between specialized and control rater pools?
  • RQ5How can identity-based rater pools improve the fairness and inclusivity of toxicity detection models?

Key findings

  • Rater identity was a statistically significant factor in toxicity annotations, with African American and LGBTQ raters assigning different ratings than the control pool for identity-related comments.
  • The identity attack, threat, and profanity categories showed the most significant differences in ratings between specialized and control rater pools.
  • Models trained on annotations from specialized rater pools outperformed models trained on larger datasets labeled by random raters, indicating higher quality and more nuanced labeling.
  • The study’s dataset of 382,500 annotations from three specialized rater pools is publicly released to support future research and model evaluation.
  • Preliminary content analysis revealed that nuanced, community-specific linguistic markers—such as microaggressions—were more accurately identified by identity-specific raters.
  • The findings suggest that involving community-identified annotators leads to more inclusive and less biased machine learning models for toxicity detection.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.