[論文レビュー] Can ChatGPT Reproduce Human-Generated Labels? A Study of Social Computing Tasks
ChatGPT は五つのソーシャル・コンピューティング・タスクで人間が生成した注釈を再現できるが、平均精度は 0.609 であり、タスクとラベルによってパフォーマンスは大きく異なる。
The release of ChatGPT has uncovered a range of possibilities whereby large language models (LLMs) can substitute human intelligence. In this paper, we seek to understand whether ChatGPT has the potential to reproduce human-generated label annotations in social computing tasks. Such an achievement could significantly reduce the cost and complexity of social computing research. As such, we use ChatGPT to relabel five seminal datasets covering stance detection (2x), sentiment analysis, hate speech, and bot detection. Our results highlight that ChatGPT does have the potential to handle these data annotation tasks, although a number of challenges remain. ChatGPT obtains an average accuracy 0.609. Performance is highest for the sentiment analysis dataset, with ChatGPT correctly annotating 64.9% of tweets. Yet, we show that performance varies substantially across individual labels. We believe this work can open up new lines of analysis and act as a basis for future research into the exploitation of ChatGPT for human annotation tasks.
研究の動機と目的
- ChatGPT がソーシャル・コンピューティングタスクにおいて人間が生成した注釈を再現できるかを評価する。
- 複数のデータセットにわたって、ChatGPT が生成したラベルを人間の真の注釈と比較する。
- タスク別およびラベル別の性能を分析し、データラベリングにおける ChatGPT の強みと限界を特定する。
- 将来の人間注釈およびクラウドソーシング文脈での大規模言語モデルの活用を導く洞察を提供する。
提案手法
- stance、ヘイトスピーチ、感情、ボット検出、ロシア・ウクライナ感情をカバーする、英語の人間注釈付き Twitter データセットを5つ選択する。
- OpenAI GPT-3.5-turbo を用いて、プロンプトテンプレートを使ってツイートを注釈する:トピックとラベルのセットが与えられたとき、ツイートを分類し説明を提供する。
- ChatGPT の主要ラベルを回答の最初の文から抽出し、全テキストを用いて説明を導出する。
- 重み付き F1 スコア、適合率、再現率を用いて人間のグラウンドトゥースと対比して ChatGPT の注釈を評価する。
実験結果
リサーチクエスチョン
- RQ1Can ChatGPT reproduce human-annotated labels on diverse social computing tasks?
- RQ2How does ChatGPT performance vary across tasks and across individual label categories?
- RQ3What are the strengths and limitations of using ChatGPT as an annotation tool in terms of precision and recall for each label?
主な発見
- Average ChatGPT annotation accuracy across five tasks is 0.609 (SD 0.032).
- Sentiment analysis yields the highest task accuracy at 0.649 with 64.9% correct labels.
- Hate speech task shows lower overall performance at 0.571, with precision issues notably for the Hate label (0.353).
- Bot detection task achieves 0.639 accuracy but exhibits large label gaps (Human 0.748 F1 vs Bot 0.364 F1).
- Russo-Ukrainian sentiment shows lowest task performance at 0.573 with notable Pro-Russia vs Pro-Ukraine label disparities (0.733 vs 0.429 F1).
- ChatGPT’s label performance varies heavily by individual labels within tasks, indicating limits for uniform annotation quality.
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。