[論文レビュー] Perception, performance, and detectability of conversational artificial intelligence across 32 university courses
本研究では、32の大学課程の授業においてChatGPTのパフォーマンスを学生の作業と比較し、2つの分類器を用いてAI生成テキストの検出可能性をテストした。ChatGPTは大多数の授業で学生を上回るか同等のパフォーマンスを示したが、現在の検出ツールは高い偽陽性率を示し、単純なテキストの難読化によって容易に回避可能であることが判明した。
The emergence of large language models has led to the development of powerful tools such as ChatGPT that can produce text indistinguishable from human-generated work. With the increasing accessibility of such technology, students across the globe may utilize it to help with their school work -- a possibility that has sparked discussions on the integrity of student evaluations in the age of artificial intelligence (AI). To date, it is unclear how such tools perform compared to students on university-level courses. Further, students' perspectives regarding the use of such tools, and educators' perspectives on treating their use as plagiarism, remain unknown. Here, we compare the performance of ChatGPT against students on 32 university-level courses. We also assess the degree to which its use can be detected by two classifiers designed specifically for this purpose. Additionally, we conduct a survey across five countries, as well as a more in-depth survey at the authors' institution, to discern students' and educators' perceptions of ChatGPT's use. We find that ChatGPT's performance is comparable, if not superior, to that of students in many courses. Moreover, current AI-text classifiers cannot reliably detect ChatGPT's use in school work, due to their propensity to classify human-written answers as AI-generated, as well as the ease with which AI-generated text can be edited to evade detection. Finally, we find an emerging consensus among students to use the tool, and among educators to treat this as plagiarism. Our findings offer insights that could guide policy discussions addressing the integration of AI into educational frameworks.
研究の動機と目的
- 多様な学術分野における大学学生の成績と比較して、ChatGPTのパフォーマンスを評価すること。
- GPTZeroとOpenAIの分類器の2つのAIテキスト分類器を用いて、ChatGPTが生成したテキストの検出可能性を評価すること。
- Quillbotなどのツールを用いた難読化攻撃が、これらの分類器に与える影響を調査すること。
- 学生および教育者らが、学術的作業におけるChatGPTの使用に関してどのように捉えているかを検討すること。
- 高等教育分野におけるAI利用に関する学術的誠実性および学生評価フレームワークの政策立案を支援すること。
提案手法
- ニューヨーク大学アブダビ校および他の機関から、32の大学レベルの授業の質問とそれに伴う学生の回答を収集した。
- 学生の提出物と直接比較できるように、すべての授業の質問に対してChatGPTを用いて回答を生成した。
- 5か国およびニューヨーク大学アブダビ校を対象に、学生および教員のAI利用に関する認識を把握するためのアンケートを実施した。
- GPTZeroとOpenAIの分類器という2つのAIテキスト分類器を用い、学生のテキストおよびChatGPTが生成したテキストを分類した。
- Quillbotを用い、複数のモードで最大の同義語置換強度を適用して、ChatGPTの出力を再表現することで難読化攻撃を実施した。
- 各分類器について正規化された混同行列を構築し、偽陽性率および偽陰性率を計算した。

実験結果
リサーチクエスチョン
- RQ1ChatGPTのパフォーマンスは、幅広い学術分野における大学課程の授業で、学生のパフォーマンスと比べてどの程度か?
- RQ2現在のAIテキスト分類器は、学術的提出物におけるChatGPT生成テキストを信頼性を持って検出できるか?
- RQ3Quillbotによる再表現といった難読化技術は、AI検出システムを回避するためにどの程度効果的か?
- RQ4学生および教育者らが、学術的作業におけるChatGPTの使用に関して、実証的および規範的認識をどのように持っているか?
- RQ5これらの発見は、高等教育分野における学術的誠実性のポリシーにどのような意味を持つのか?
主な発見
- ChatGPTは32の大学課程の授業のうち28で学生の成績を上回ったか同等の成績を示し、特に人文学および社会科学分野で顕著な結果を示した。
- GPTZero分類器は、32.5%の人の書いた学生の回答をAI生成と誤って分類しており、高い偽陽性率を示している。
- OpenAIの分類器は、28.3%の人の回答をAI生成と誤って分類しており、検出システムの信頼性の低さをさらに浮き彫りにしている。
- GPTZeroは12.5%のChatGPT出力を人間が書いたと誤って分類し、OpenAIの分類器は15.7%を同様に誤分類しており、非軽微な偽陰性率があることが示された。
- Quillbotを用いた難読化攻撃は、両方の分類器において68.8%のChatGPT出力を検出回避に成功しており、高い回避成功率を示した。
- アンケート調査において、教育者の87%がAI生成の作業を不正使用とみなした一方、学生の73%が学術的タスクでChatGPTを利用していた。

より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。