Skip to main content
QUICK REVIEW

[论文解读] Is ChatGPT better than Human Annotators? Potential and Limitations of ChatGPT in Explaining Implicit Hate Speech

Huang Fan, Haewoon Kwak|arXiv (Cornell University)|Feb 11, 2023
Hate Speech and Cyberbullying Detection被引用 28
一句话总结

论文评估 ChatGPT 在检测隐含仇恨言论和生成自然语言解释方面的能力,发现与人工标注数据集有 80% 的一致性,且 ChatGPT 的解释比人类更清晰,信息性相当。

ABSTRACT

Recent studies have alarmed that many online hate speeches are implicit. With its subtle nature, the explainability of the detection of such hateful speech has been a challenging problem. In this work, we examine whether ChatGPT can be used for providing natural language explanations (NLEs) for implicit hateful speech detection. We design our prompt to elicit concise ChatGPT-generated NLEs and conduct user studies to evaluate their qualities by comparison with human-written NLEs. We discuss the potential and limitations of ChatGPT in the context of implicit hateful speech research.

研究动机与目标

  • 评估 ChatGPT 是否能像人类一样有效地检测隐含仇恨推文。
  • 评估 ChatGPT 生成的隐含仇恨言论的自然语言解释的质量。
  • 在信息性和清晰度方面,将 ChatGPT 的解释与人类撰写的解释进行比较。

提出的方法

  • 使用 6,358 条隐含仇恨推文的 LatentHatred 数据集作为测试床。
  • 对每条推文生成三份 ChatGPT 回应,使用特定提示并将 +1/0/-1 的平均分作为 ChatGPT 仇恨分数。
  • 进行 MTurk 基于的人类评估,以在三种情境下比较 ChatGPT 的分类和 NLEs:仅贴文、贴文+人类 NLE、贴文+ChatGPT NLE。
  • 使用七点量表的 信息性 与 清晰度 指标来评估 NLE 的质量。
  • 将 ChatGPT 结果与人类标注进行比较,以评估一致性和潜在偏差。

实验结果

研究问题

  • RQ1RQ1:ChatGPT 能否很好地检测隐含仇恨推文?
  • RQ2RQ2:ChatGPT 是否为隐含仇恨言论生成高质量的自然语言解释?

主要发现

  • ChatGPT 正确识别了 795 条隐含仇恨推文中的 636 条(与 LatentHatred 标签的 80% 一致)。
  • 在仅给定贴文的情况下,ChatGPT 产生的仇恨性平均分为 -0.41;在提供 ChatGPT 生成的 NLE 时为 -0.52,表明普通人对推文的判断往往认为其不具仇恨性,且 NLE 影响判断。
  • 提供人类撰写的 NLE 可得到正的平均分 0.29,显示人类解释可以改变判断,但总体上不如 ChatGPT 的解释有说服力。
  • ChatGPT 生成的 NLE 在清晰度上更高(均值 5.39)低于人类撰写的 NLE(均值 4.68),信息性方面没有显著差异。
  • ChatGPT 的表现表明在主观任务上作为数据标注工具具有很大潜力,尽管如果其决策有误也存在风险。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。