[논문 리뷰] How would Stance Detection Techniques Evolve after the Launch of ChatGPT?
논문은 SemEval-2016 및 P-Stance 벤치마크에서 ChatGPT의 제로샷 및 도메인 내 성능을 조사하여, ChatGPT가 SOTA 또는 근접 SOTA 결과에 도달하고 예측에 대한 설명을 제공할 수 있어 stance 탐지 연구 패러다임을 바꿀 수 있음을 시사한다.
Stance detection refers to the task of extracting the standpoint (Favor, Against or Neither) towards a target in given texts. Such research gains increasing attention with the proliferation of social media contents. The conventional framework of handling stance detection is converting it into text classification tasks. Deep learning models have already replaced rule-based models and traditional machine learning models in solving such problems. Current deep neural networks are facing two main challenges which are insufficient labeled data and information in social media posts and the unexplainable nature of deep learning models. A new pre-trained language model chatGPT was launched on Nov 30, 2022. For the stance detection tasks, our experiments show that ChatGPT can achieve SOTA or similar performance for commonly used datasets including SemEval-2016 and P-Stance. At the same time, ChatGPT can provide explanation for its own prediction, which is beyond the capability of any existing model. The explanations for the cases it cannot provide classification results are especially useful. ChatGPT has the potential to be the best AI model for stance detection tasks in NLP, or at least change the research paradigm of this field. ChatGPT also opens up the possibility of building explanatory AI for stance detection.
연구 동기 및 목표
- 표준 데이터셋(SemEval-2016, P-Stance)에서 ChatGPT의 제로샷 stance 탐지 성능 평가.
- 데이터에 대해 학습된 전통적 모델과 비교한 도메인 내 ChatGPT 성능 평가.
- ChatGPT의 stance 결정에 대한 설명 능력 및 stance 탐지를 위한 설명 가능한 AI 가능성 분석.
제안 방법
- 제로샷 프롬프팅으로 주어진 트윗의 대상에 대한 입장을 ChatGPT에 직접 요청하는 프롬프트 구성.
- SemEval-2016 및 P-Stance 데이터셋에서 F1-avg 및 macro-F1 (F1-m)으로 출력 평가.
- 제로샷 ChatGPT 결과를 전통적 모델 및 도메인 내(80% 학습 데이터) 베이스라인과 비교.
- ChatGPT가 입장을 제공할 수 없을 때, 그 설명을 분석하여 향후 설명 가능한 AI 인사이트를 도출.
- 논문의 업데이트 버전에서 GPT-3.5-0301 변형으로 결과를 업데이트.
실험 결과
연구 질문
- RQ1표준 데이터셋에서 제로샷 설정에서 ChatGPT가 최첨단 또는 경쟁력 있는 stance 탐지 성능을 달성할 수 있는가?
- RQ2도메인 내에서 상당량의 라벨링 데이터를 학습한 모델과 비교했을 때 ChatGPT의 성능은 어떠한가?
- RQ3ChatGPT가 자신의 stance 결정에 대해 어떤 종류의 설명을 제공하며, 다턴 프롬프팅이 성능을 더 향상시킬 수 있는가?
- RQ4설명 가능한 AI를 위한 프롬프트 설계, 설명 가능성, 다라운드 대화를 통해 stance 탐지에서 어떤 미래 연구 방향이 도출되는가?
주요 결과
- 제로샷 프롬프팅 하에서 SemEval-2016 및 P-Stance에서 ChatGPT가 최첨단 또는 유사 한 성능을 달성.
- 제로샷 설정에서 ChatGPT가 종종 베이스라인을 능가하고 도메인 내 설정에서도 경쟁력을 유지.
- ChatGPT는 명시적 및 암시적 stance 신호를 포함한 의사 결정에 대한 설명을 제공할 수 있다.
- 다단계 프롬프트(Chain-of-prompts, 다턴 상호작용)가 stance 탐지 성능을 더 개선할 수 있다.
- 향후 세 가지 제시 방향으로 더 나은 프롬프트 템플릿, 설명 가능한 AI를 위한 ChatGPT의 설명 활용, 어려운 사례에 대한 커버리지를 개선하기 위한 다라운드 대화 탐색이 제안된다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.