[论文解读] "HOT" ChatGPT: The promise of ChatGPT in detecting and discriminating hateful, offensive, and toxic comments on social media
本论文评估 ChatGPT 发现令人讨厌、冒犯性和有害(HOT)评论的能力,并将其性能与 MTurk 注释在五个提示和四个实验中进行比较,结果约为 80% 准确率,并强调提示敏感性及与 HOT 定义的一致性。
Harmful content is pervasive on social media, poisoning online communities and negatively impacting participation. A common approach to address this issue is to develop detection models that rely on human annotations. However, the tasks required to build such models expose annotators to harmful and offensive content and may require significant time and cost to complete. Generative AI models have the potential to understand and detect harmful content. To investigate this potential, we used ChatGPT and compared its performance with MTurker annotations for three frequently discussed concepts related to harmful content: Hateful, Offensive, and Toxic (HOT). We designed five prompts to interact with ChatGPT and conducted four experiments eliciting HOT classifications. Our results show that ChatGPT can achieve an accuracy of approximately 80% when compared to MTurker annotations. Specifically, the model displays a more consistent classification for non-HOT comments than HOT comments compared to human annotations. Our findings also suggest that ChatGPT classifications align with provided HOT definitions, but ChatGPT classifies "hateful" and "offensive" as subsets of "toxic." Moreover, the choice of prompts used to interact with ChatGPT impacts its performance. Based on these in-sights, our study provides several meaningful implications for employing ChatGPT to detect HOT content, particularly regarding the reliability and consistency of its performance, its understand-ing and reasoning of the HOT concept, and the impact of prompts on its performance. Overall, our study provides guidance about the potential of using generative AI models to moderate large volumes of user-generated content on social media.
研究动机与目标
- 在不需要人类注释者接触有害材料的前提下,推动使用生成式人工智能来 Moderating 大量用户生成内容。
- 研究 ChatGPT 分类 HOT 内容的能力,并与基于标准 HOT 定义的 MTurk 注释进行比较。
- 考察不同提示对 ChatGPT 的表现及其与 HOT 概念(仇恨、冒犯、有毒)的对齐程度的影响。
- 就 ChatGPT 在 HOT 内容检测中的可靠性、一致性和推理能力提供指南。
提出的方法
- 设计五个提示与 ChatGPT 进行 HOT 分类交互。
- 进行四个实验,从 ChatGPT 获取 HOT 分类结果。
- 比较 ChatGPT 的分类与对仇恨、冒犯和有害内容的 MTurk 注释。
- 分析 ChatGPT 对 HOT 与非 HOT 评论分类的一致性。
- 考察 ChatGPT 将仇恨和冒犯视为有害的子集的情况,以及提示如何影响结果。
实验结果
研究问题
- RQ1ChatGPT 是否能像 MTurk 注释那样准确地检测和区分 HOT 内容?
- RQ2ChatGPT 对 HOT 与非 HOT 评论的 HOT 分类有多一致?
- RQ3ChatGPT 的分类是否与提供的 HOT 定义一致,提示如何影响这种对齐?
- RQ4ChatGPT 将仇恨和冒犯视为有害的子集吗?这对 moderation 有何启示?
- RQ5提示选择对 ChatGPT 在 HOT 检测中的表现有何影响?
主要发现
- ChatGPT 相对于 MTurk 注释的准确率大约为 80%。
- 相较于人工注释,ChatGPT 对非 HOT 评论的分类更具一致性,而对 HOT 评论的一致性较低。
- ChatGPT 的分类与提供的 HOT 定义一致。
- ChatGPT 往往将仇恨和冒犯视为有害的子集。
- 用于与 ChatGPT 交互的提示选择会影响其性能。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。