Skip to main content
QUICK REVIEW

[Paper Review] Is ChatGPT better than Human Annotators? Potential and Limitations of ChatGPT in Explaining Implicit Hate Speech

Huang Fan, Haewoon Kwak|arXiv (Cornell University)|Feb 11, 2023
Hate Speech and Cyberbullying Detection28 citations
TL;DR

The paper evaluates ChatGPT for detecting implicit hate speech and generating natural language explanations, finding 80% agreement with a human-annotated dataset and that ChatGPT explanations are clearer than human ones, with comparable informativeness.

ABSTRACT

Recent studies have alarmed that many online hate speeches are implicit. With its subtle nature, the explainability of the detection of such hateful speech has been a challenging problem. In this work, we examine whether ChatGPT can be used for providing natural language explanations (NLEs) for implicit hateful speech detection. We design our prompt to elicit concise ChatGPT-generated NLEs and conduct user studies to evaluate their qualities by comparison with human-written NLEs. We discuss the potential and limitations of ChatGPT in the context of implicit hateful speech research.

Motivation & Objective

  • Assess whether ChatGPT can detect implicit hateful tweets as effectively as humans.
  • Evaluate the quality of ChatGPT-generated natural language explanations for implicit hate speech.
  • Compare ChatGPT explanations with human-written explanations in terms of informativeness and clarity.

Proposed method

  • Use LatentHatred dataset of 6,358 implicit hateful tweets as the testbed.
  • Generate three ChatGPT responses per tweet with a specific prompt and average a +1/0/-1 scoring to derive a ChatGPT hate score.
  • Conduct MTurk-based human evaluations to compare ChatGPT classification and NLEs under three contexts: post only, post+human NLE, and post+ChatGPT NLE.
  • Assess NLE quality using Informativeness and Clarity metrics on a 7-point scale.
  • Compare ChatGPT results with human annotations to assess agreement and potential biases.

Experimental results

Research questions

  • RQ1RQ1: Can ChatGPT detect implicit hateful tweets well?
  • RQ2RQ2: Does ChatGPT generate quality natural language explanations for implicit hate speech?

Key findings

  • ChatGPT correctly identifies 636 of 795 implicit hateful tweets (80% agreement with the LatentHatred labels).
  • ChatGPT yields -0.41 average hatefulness score when given a post alone and -0.52 when provided with ChatGPT-generated NLEs, indicating laypeople’reactions tend to deem tweets non-hateful, and the NLEs influence judgments.
  • Providing human-written NLEs leads to a positive average score of 0.29, showing human explanations can shift judgments but are less convincing overall than ChatGPT’s explanations.
  • ChatGPT-generated NLEs have higher clarity (mean 5.39) than human-written NLEs (mean 4.68), with no significant difference in informativeness.
  • ChatGPT’s performance suggests strong potential as a data-annotation tool for subjective tasks, though risks exist if its decisions are wrong.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.