[논문 리뷰] "HOT" ChatGPT: The promise of ChatGPT in detecting and discriminating hateful, offensive, and toxic comments on social media
이 논문은 ChatGPT의 혐오적, 모욕적, 독성(HOT) 코멘트 탐지 능력을 평가하고 다섯 가지 프롬프트와 네 가지 실험에서 MTurk 주석과의 성능을 비교하며 약 80% 정확도와 프롬프트 민감도 및 HOT 정의와의 정렬을 강조합니다.
Harmful content is pervasive on social media, poisoning online communities and negatively impacting participation. A common approach to address this issue is to develop detection models that rely on human annotations. However, the tasks required to build such models expose annotators to harmful and offensive content and may require significant time and cost to complete. Generative AI models have the potential to understand and detect harmful content. To investigate this potential, we used ChatGPT and compared its performance with MTurker annotations for three frequently discussed concepts related to harmful content: Hateful, Offensive, and Toxic (HOT). We designed five prompts to interact with ChatGPT and conducted four experiments eliciting HOT classifications. Our results show that ChatGPT can achieve an accuracy of approximately 80% when compared to MTurker annotations. Specifically, the model displays a more consistent classification for non-HOT comments than HOT comments compared to human annotations. Our findings also suggest that ChatGPT classifications align with provided HOT definitions, but ChatGPT classifies "hateful" and "offensive" as subsets of "toxic." Moreover, the choice of prompts used to interact with ChatGPT impacts its performance. Based on these in-sights, our study provides several meaningful implications for employing ChatGPT to detect HOT content, particularly regarding the reliability and consistency of its performance, its understand-ing and reasoning of the HOT concept, and the impact of prompts on its performance. Overall, our study provides guidance about the potential of using generative AI models to moderate large volumes of user-generated content on social media.
연구 동기 및 목표
- 대용량의 사용자 생성 콘텐츠를 인간 주석가의 노출 없이 관리하기 위해 생성형 AI의 사용을 촉진한다.
- ChatGPT의 HOT 콘텐츠 분류 능력을 조사하고 표준 HOT 정의에 비춰 MTurk 주석과 비교한다.
- 다양한 프롬프트가 ChatGPT의 성능 및 HOT 개념(혐오, 모욕, 독성)과의 정렬에 어떤 영향을 미치는지 살펴본다.
- HOT 콘텐츠 탐지에서 ChatGPT의 신뢰성, 일관성 및 추론에 대한 가이드를 제공한다.
제안 방법
- HOT 분류를 위해 ChatGPT와 상호작용하도록 다섯 가지 프롬프트를 설계한다.
- ChatGPT로부터 HOT 분류를 이끌어내는 네 가지 실험을 수행한다.
- 혐오, 모욕, 독성 콘텐츠에 대한 ChatGPT 분류를 MTurk 주석과 비교한다.
- HOT 대 비HOT 코멘트에 대한 ChatGPT의 분류 일관성을 분석한다.
- ChatGPT가 혐오와 모욕을 독성의 하위 집합으로 취급하는지 여부와 프롬프트가 결과에 미치는 영향을 검토한다.
실험 결과
연구 질문
- RQ1ChatGPT가 MTurk 주석과 비교해 HOT 내용을 정확하게 탐지하고 구분할 수 있는가?
- RQ2HOT 및 비HOT 코멘트에 대한 ChatGPT의 HOT 분류 일관성은 얼마나 되는가?
- RQ3ChatGPT 분류가 제공된 HOT 정의와 얼마나 일치하며 프롬프트가 이 정렬에 어떤 영향을 주는가?
- RQ4혐오와 모욕이 ChatGPT에 의해 독성의 하위 집합으로 간주되는가, 그리고 이것이 중재에 어떤 시사점을 주는가?
- RQ5HOT 탐지에서 ChatGPT의 성능에 프롬프트 선택이 어떤 영향을 미치는가?
주요 결과
- ChatGPT는 MTurk 주석에 비해 약 80%의 정확도를 달성한다.
- ChatGPT는 인간 주석에 비해 HOT가 아닌 코멘트에 대해 더 일관된 분류를 보이는 반면 HOT 코멘트에 대해서는 그렇지 않다.
- ChatGPT 분류는 제공된 HOT 정의와 일치한다.
- ChatGPT는 혐오와 모욕을 독성의 하위 집합으로 분류하는 경향이 있다.
- ChatGPT와 상호작용하는 데 사용되는 프롬프트의 선택이 성능에 영향을 미친다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.