[Paper Review] "HOT" ChatGPT: The promise of ChatGPT in detecting and discriminating hateful, offensive, and toxic comments on social media
The paper evaluates ChatGPT’s ability to detect hateful, offensive, and toxic (HOT) comments and compares its performance to MTurk annotations across five prompts and four experiments, finding ~80% accuracy and highlighting prompt sensitivity and alignment with HOT definitions.
Harmful content is pervasive on social media, poisoning online communities and negatively impacting participation. A common approach to address this issue is to develop detection models that rely on human annotations. However, the tasks required to build such models expose annotators to harmful and offensive content and may require significant time and cost to complete. Generative AI models have the potential to understand and detect harmful content. To investigate this potential, we used ChatGPT and compared its performance with MTurker annotations for three frequently discussed concepts related to harmful content: Hateful, Offensive, and Toxic (HOT). We designed five prompts to interact with ChatGPT and conducted four experiments eliciting HOT classifications. Our results show that ChatGPT can achieve an accuracy of approximately 80% when compared to MTurker annotations. Specifically, the model displays a more consistent classification for non-HOT comments than HOT comments compared to human annotations. Our findings also suggest that ChatGPT classifications align with provided HOT definitions, but ChatGPT classifies "hateful" and "offensive" as subsets of "toxic." Moreover, the choice of prompts used to interact with ChatGPT impacts its performance. Based on these in-sights, our study provides several meaningful implications for employing ChatGPT to detect HOT content, particularly regarding the reliability and consistency of its performance, its understand-ing and reasoning of the HOT concept, and the impact of prompts on its performance. Overall, our study provides guidance about the potential of using generative AI models to moderate large volumes of user-generated content on social media.
Motivation & Objective
- Motivate the use of generative AI for moderating large volumes of user-generated content without requiring human annotators exposure to harmful material.
- Investigate ChatGPT’s capability to classify HOT content and compare it with MTurk annotations across standard HOT definitions.
- Examine how different prompts influence ChatGPT’s performance and alignment with HOT concepts (hateful, offensive, toxic).
- Provide guidance on the reliability, consistency, and reasoning of ChatGPT in HOT content detection.
Proposed method
- Design five prompts to interact with ChatGPT for HOT classification.
- Conduct four experiments eliciting HOT classifications from ChatGPT.
- Compare ChatGPT classifications with MTurk annotations on hateful, offensive, and toxic content.
- Analyze consistency of ChatGPT’s classifications for HOT vs non-HOT comments.
- Examine whether ChatGPT treats hateful and offensive as subsets of toxic and how prompts affect results.
Experimental results
Research questions
- RQ1Can ChatGPT accurately detect and discriminate HOT content compared with MTurk annotations?
- RQ2How consistent are ChatGPT’s HOT classifications for HOT versus non-HOT comments?
- RQ3Do ChatGPT classifications align with provided HOT definitions, and how do prompts influence this alignment?
- RQ4Are hateful and offensive considered subsets of toxic by ChatGPT, and what does this imply for moderation?
- RQ5What is the impact of prompt choice on ChatGPT’s performance in HOT detection?
Key findings
- ChatGPT achieves approximately 80% accuracy compared to MTurk annotations.
- ChatGPT shows more consistent classification for non-HOT comments than for HOT comments relative to human annotations.
- ChatGPT classifications align with the provided HOT definitions.
- ChatGPT tends to classify hateful and offensive as subsets of toxic.
- The choice of prompts used to interact with ChatGPT impacts its performance.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.