[論文レビュー] Comparing Abstractive Summaries Generated by ChatGPT to Real Summaries Through Blinded Reviewers and Text Classification Algorithms
本論文は、ChatGPT が生成した抽象的要約を実際の要約と自動指標、ブラインド評価、人間の評価、テキスト分類器を用いて比較し、人間は二者を区別するのが困難である一方、分類器は高い精度で識別できることを示している。
Large Language Models (LLMs) have gathered significant attention due to their impressive performance on a variety of tasks. ChatGPT, developed by OpenAI, is a recent addition to the family of language models and is being called a disruptive technology by a few, owing to its human-like text-generation capabilities. Although, many anecdotal examples across the internet have evaluated ChatGPT's strength and weakness, only a few systematic research studies exist. To contribute to the body of literature of systematic research on ChatGPT, we evaluate the performance of ChatGPT on Abstractive Summarization by the means of automated metrics and blinded human reviewers. We also build automatic text classifiers to detect ChatGPT generated summaries. We found that while text classification algorithms can distinguish between real and generated summaries, humans are unable to distinguish between real summaries and those produced by ChatGPT.
研究の動機と目的
- ROUGE、METEOR などの自動指標を用いて、ChatGPT の要約生成(抽象的要約)を実際の要約と比較評価する。
- ブラインド評価を行う人間のレビュアーが、ChatGPT が生成した要約と実際の要約を識別できるかを評価する。
- ChatGPT が生成した要約を検出するためのテキスト分類器を開発・評価する。
提案手法
- Nallapati ら (2016) の CNN/Daily News 要約 50 件のデータセットを作成する。
- 同じ 50 記事について、慎重に選択したプロンプトで ChatGPT に要約を生成させる。
- 実際の要約と生成した要約の間で自動指標(ROUGE-1/2/L、ROUGE-LSUM、METEOR)を計算する。
- 50 ペアで、母語話者の英語話者2名を用いたブラインド評価を実施する。
- ChatGPT 対 人間の要約検出のため DistillBERT を微調整し、Sentence Embeddings + XGBoost と比較する。

実験結果
リサーチクエスチョン
- RQ1ChatGPT は、人間には実際の要約と区別できない抽象的要約を生成できるか?
- RQ2自動指標は、実際の要約と ChatGPT が生成した要約の差を定量化できるか?
- RQ3分類器は、ChatGPT生成と人間作成の要約を信頼性高く区別できるか?
主な発見
- 自動指標は、実際の要約と ChatGPT が生成した要約の間に測定可能な差を示す(ROUGE-1: 0.30、ROUGE-2: 0.11、ROUGE-L: 0.20、ROUGELSUM: 0.21、METEOR: 0.35)。
- ブラインド評価の人間レビュアーは、生成と実際の要約を区別できず、正確さは約 0.49 だった。
- 微調整した DistillBERT は、ChatGPT 生成と実際の要約を 90% の精度(F1 0.33)で検出する。
- Sentence Embeddings と XGBoost による 50% の精度(ベースライン)を達成。
- 本研究は、プロンプトを慎重に選択すれば、人間は ChatGPT の要約を実際のものと同等に認識する傾向があることを示唆している。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。