Skip to main content
QUICK REVIEW

[논문 리뷰] ChatGPT: Beginning of an End of Manual Linguistic Data Annotation? Use Case of Automatic Genre Identification

Taja Kuzman, Igor Mozetič|arXiv (Cornell University)|2023. 03. 07.
Text Readability and Simplification인용 수 66
한 줄 요약

본 논문은 자동 장르 식별을 위한 제로샷 ChatGPT를 평가하고 영어와 슬로베니아어에서 미세 조정된 X-GENRE 모델과 비교하며, ChatGPT가 보지 못한 데이터에서 미세 조정 모델을 능가할 수 있고 프롬프트 언어가 성능에 영향을 준다는 것을 보여준다.

ABSTRACT

ChatGPT has shown strong capabilities in natural language generation tasks, which naturally leads researchers to explore where its abilities end. In this paper, we examine whether ChatGPT can be used for zero-shot text classification, more specifically, automatic genre identification. We compare ChatGPT with a multilingual XLM-RoBERTa language model that was fine-tuned on datasets, manually annotated with genres. The models are compared on test sets in two languages: English and Slovenian. Results show that ChatGPT outperforms the fine-tuned model when applied to the dataset which was not seen before by either of the models. Even when applied on Slovenian language as an under-resourced language, ChatGPT's performance is no worse than when applied to English. However, if the model is fully prompted in Slovenian, the performance drops significantly, showing the current limitations of ChatGPT usage on smaller languages. The presented results lead us to questioning whether this is the beginning of an end of laborious manual annotation campaigns even for smaller languages, such as Slovenian.

연구 동기 및 목표

  • 영어와 슬로베니아어에서 자동 장르 식별에 대한 ChatGPT의 제로샷 성능 평가.
  • 다국어 장르 데이터세트로 학습된 미세 조정된 X-GENRE 분류기와의 비교.
  • 프롬프트 언어가 ChatGPT의 분류 성능에 미치는 영향 조사.
  • ChatGPT와 X-GENRE 결과 간의 모델 일치도 및 보완성 분석.
  • NLP 연구에서 수작업 주석 노력에 미치는 시사점 논의.

제안 방법

  • EN-GINCO(English)와 GINCO(Slovenian) 데이터세트를 사용하며, 각 데이터세트는 100개의 테스트 인스턴스로 X-GENRE 스키마로 라벨링된다.
  • CORE, FTD, GINCO 데이터세트의 약 1,700개 수동 주석 텍스트에 대해 X-GENRE(XLM-RoBERTa)를 미세 조정하고 X-GENRE 라벨에 매핑한다.
  • 세 가지 시나리오에서 ChatGPT에 프롬프트를 제공한다: 영어 텍스트에 영어 프롬프트, 슬로베니아어 텍스트에 영어 프롬프트, 슬로베니아어 텍스트에 슬로베니아어 프롬프트.
  • 출력에서 ChatGPT 예측과 설명을 추출하고 X-GENRE 스키마의 실제 라벨과 비교 평가한다.
  • X-GENRE에 비해 마이크로 F1, 매크로 F1, 정확도를 비교하고 불일치를 분석하여 보완 강점을 평가한다.
Figure 1: Comparison of differences in correct and incorrect predictions between ChatGPT and X-GENRE.
Figure 1: Comparison of differences in correct and incorrect predictions between ChatGPT and X-GENRE.

실험 결과

연구 질문

  • RQ1보지 못한 데이터에서 제로샷 장르 식별을 미세 조정된 다국어 분류기와 비교해 유사하게 수행할 수 있는가?
  • RQ2영어와 슬로베니아어 텍스트, 서로 다른 언어로 프롬프트를 제공했을 때 ChatGPT의 성능 차이는 어떻게 나타나는가?
  • RQ3프롬프트 언어가 ChatGPT의 장르 분류 정확도에 미치는 영향은 무엇인가?
  • RQ4ChatGPT와 X-GENRE가 함께 사용될 때 서로 보완하는 예측을 만들어 성능 향상에 기여할 수 있는가?

주요 결과

  • ChatGPT가 영어 EN-GINCO 테스트 세트에서 영어 프롬프트일 때 X-GENRE를 능가한다(micro F1 0.74, macro F1 0.66, accuracy 0.72 대 0.67/0.61/0.67).
  • 슬로베니아 데이터(GINCO)에서 슬로베니아어 텍스트로 슬로베니아어 프롬프트를 사용할 경우 X-GENRE가 ChatGPT를 상당히 능가한다(micro F1 0.91, macro F1 0.91, accuracy 0.91 for X-GENRE; ChatGPT 0.68/0.56/0.68 for Slovenian prompt).
  • 영어 프롬프트일 때 슬로베니아어 텍스트에 대한 ChatGPT의 성능은 영어와 비교할 만하며, 슬로베니아어는 자원이 부족한 점에도 불구하고 영어 프롬프트에서의 성능이 더 좋고, 슬로베니아어 프롬프트일 때 성능이 저하된다.
  • ChatGPT의 예측은 X-GENRE와 대다수에서 보완적이며, 서로 다른 다수 오류를 보여 특정 사용 사례에서 앙상블 이득의 가능성을 시사한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.