[Paper Review] Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive
This paper introduces a large-scale noise audit and the novel concept of vicarious offense to study disagreement in offensive speech classification across human and machine moderators in political discourse. It reveals extensive variability in moderation outcomes and shows that neither humans nor large language models can reliably predict how others—especially those with differing political leanings—will perceive offensive content.
Offensive speech detection is a key component of content moderation. However, what is offensive can be highly subjective. This paper investigates how machine and human moderators disagree on what is offensive when it comes to real-world social web political discourse. We show that (1) there is extensive disagreement among the moderators (humans and machines); and (2) human and large-language-model classifiers are unable to predict how other human raters will respond, based on their political leanings. For (1), we conduct a noise audit at an unprecedented scale that combines both machine and human responses. For (2), we introduce a first-of-its-kind dataset of vicarious offense. Our noise audit reveals that moderation outcomes vary wildly across different machine moderators. Our experiments with human moderators suggest that political leanings combined with sensitive issues affect both first-person and vicarious offense. The dataset is available through https://github.com/Homan-Lab/voiced.
Motivation & Objective
- To investigate the extent of disagreement among human and machine moderators in identifying offensive political discourse on social media.
- To address the lack of large-scale, real-world evaluations of offensive speech classifiers beyond controlled datasets.
- To introduce the concept of vicarious offense—how individuals perceive whether content would offend others with different political identities.
- To evaluate whether political leanings influence perceptions of both first-person and vicarious offense.
- To assess the ability of large language models and traditional classifiers to predict human disagreement in offensive speech judgments.
Proposed method
- Conduct a noise audit using 9 machine moderators (including LLMs like GPT-3.5 and non-LLM models) and human raters across a large-scale YouTube comment dataset with known political dissonance.
- Collect human annotations on both first-person (personal) and vicarious offense perception using a structured prompt format: 'How offensive do you think [partyA] will find this comment?'
- Use majority voting to aggregate machine and human moderator judgments, enabling comparison across systems and political groups.
- Introduce the VOICED dataset—a first-of-its-kind collection of human-labeled vicarious offense annotations across Democratic, Republican, and Independent perspectives.
- Evaluate model performance using metrics like F1, precision, and recall, comparing LLMs (e.g., ChatGPT) and non-LLM classifiers (e.g., Perspective API, Detoxify) against human consensus.
- Apply the noise audit framework inspired by Kahneman et al. to measure outcome variability across competent decision systems, treating each moderator as a distinct decision-maker.

Experimental results
Research questions
- RQ1How much disagreement exists among human and machine moderators in labeling offensive political discourse?
- RQ2To what extent do political leanings influence perceptions of first-person and vicarious offense?
- RQ3Can large language models accurately predict how others—especially those with opposing political views—will judge offensive content?
- RQ4How does the performance of machine moderators compare to human moderators in identifying both personal and vicarious offense?
- RQ5Can the concept of vicarious offense help unify understanding of human and machine disagreement in offensive speech detection?
Key findings
- There is extensive disagreement among both human and machine moderators in labeling offensive content, with high variability in moderation outcomes even across well-known classifiers.
- Machine moderators (including GPT-3.5 and non-LLM models) show significant inconsistency, with no single model achieving high agreement with human consensus.
- Human moderators with different political leanings disagree not only on personal offense but also on vicarious offense, indicating that political polarization affects perceived offensiveness beyond personal experience.
- Large language models like ChatGPT show limited ability to predict vicarious offense, agreeing less with human raters and more with the machine moderator majority, highlighting challenges in modeling subjective perception.
- The study reveals that political identity significantly affects both first-person and vicarious offense perception, with Independents often serving as a middle ground, while Democrats and Republicans diverge in their judgments.
- The VOICED dataset, collected via a controlled annotation pipeline, provides a new benchmark for studying vicarious offense and is publicly available for future research.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.