Skip to main content
QUICK REVIEW

[논문 리뷰] Comparing Abstractive Summaries Generated by ChatGPT to Real Summaries Through Blinded Reviewers and Text Classification Algorithms

Mayank Soni, Vincent Wade|arXiv (Cornell University)|2023. 03. 30.
Topic Modeling인용 수 9
한 줄 요약

이 연구는 ChatGPT가 생성한 추상 요약과 실제 요약을 자동 지표, 블라인드 인간 평가, 그리고 텍스트 분류기를 사용해 비교합니다. 인간은 둘을 구분하기 어렵지만 분류기는 높은 정확도로 구분할 수 있음을 발견했습니다.

ABSTRACT

Large Language Models (LLMs) have gathered significant attention due to their impressive performance on a variety of tasks. ChatGPT, developed by OpenAI, is a recent addition to the family of language models and is being called a disruptive technology by a few, owing to its human-like text-generation capabilities. Although, many anecdotal examples across the internet have evaluated ChatGPT's strength and weakness, only a few systematic research studies exist. To contribute to the body of literature of systematic research on ChatGPT, we evaluate the performance of ChatGPT on Abstractive Summarization by the means of automated metrics and blinded human reviewers. We also build automatic text classifiers to detect ChatGPT generated summaries. We found that while text classification algorithms can distinguish between real and generated summaries, humans are unable to distinguish between real summaries and those produced by ChatGPT.

연구 동기 및 목표

  • 자동 지표(ROUGE, METEOR)를 사용하여 ChatGPT의 추상 요약을 실제 요약과 비교 평가한다.
  • 블라인드 처리된 인간 평가자가 ChatGPT가 생성한 요약과 실제 요약을 구분할 수 있는지 평가한다.
  • ChatGPT가 생성한 요약을 탐지하기 위한 텍스트 분류기를 개발하고 평가한다.

제안 방법

  • Nallapati 등(2016)의 50개의 CNN/Daily News 요약 데이터셋을 생성한다.
  • 같은 50개 기사에 대해 추상 요약을 생성하기 위해 신중하게 선택된 프롬프트로 ChatGPT에 프롬프트를 제공한다.
  • 실제 요약과 생성 요약 사이의 자동 지표(ROUGE-1/2/L, ROUGE-LSUM, METEOR)를 계산한다.
  • 50쌍에 대해 두 명의 영어 원어민 평가자와 함께 블라인드 인간 평가를 수행한다.
  • ChatGPT 대비 인간 요약 감지를 위해 DistillBERT를 미세조정하고 Sentence Embeddings + XGBoost와 비교한다.
Figure 1: User Interface of Summaries Shown to Human Reviewers
Figure 1: User Interface of Summaries Shown to Human Reviewers

실험 결과

연구 질문

  • RQ1ChatGPT가 인간에게 실제 요약과 구분되지 않는 추상적 요약을 생성할 수 있는가?
  • RQ2자동 지표가 실제 요약과 ChatGPT가 생성한 요약 간의 차이를 정량화할 수 있는가?
  • RQ3분류기가 ChatGPT가 생성한 요약과 인간이 작성한 요약을 신뢰성 있게 구분할 수 있는가?

주요 결과

  • 자동 지표는 실제 요약과 ChatGPT가 생성한 요약 간에 측정 가능한 차이가 있음을 보여준다(ROUGE-1: 0.30, ROUGE-2: 0.11, ROUGE-L: 0.20, ROUGELSUM: 0.21, METEOR: 0.35).
  • 블라인드 처리된 인간 평가자는 생성 요약과 실제 요약을 구분하는 데 정확도 약 0.49로 실패했다.
  • 미세조정된 DistillBERT는 ChatGPT가 생성한 요약과 실제 요약을 90% 정확도(F1 0.33)로 감지한다.
  • Sentence Embeddings와 XGBoost는 50% 정확도(베이스라인)를 달성했다.
  • 본 연구는 프롬프트를 신중하게 선택하면 인간이 ChatGPT 요약을 실제 요약과 비교 가능하다고 인식할 수 있음을 시사한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.