[论文解读] Perception, performance, and detectability of conversational artificial intelligence across 32 university courses
本研究评估了ChatGPT在32门大学课程中的表现,将其与学生作业进行对比,并测试了两种分类器对AI生成文本的检测能力。研究发现,ChatGPT在大多数课程中表现优于或等同于学生,而当前的检测工具存在较高的误报率,且极易通过简单的文本混淆技术规避。
The emergence of large language models has led to the development of powerful tools such as ChatGPT that can produce text indistinguishable from human-generated work. With the increasing accessibility of such technology, students across the globe may utilize it to help with their school work -- a possibility that has sparked discussions on the integrity of student evaluations in the age of artificial intelligence (AI). To date, it is unclear how such tools perform compared to students on university-level courses. Further, students' perspectives regarding the use of such tools, and educators' perspectives on treating their use as plagiarism, remain unknown. Here, we compare the performance of ChatGPT against students on 32 university-level courses. We also assess the degree to which its use can be detected by two classifiers designed specifically for this purpose. Additionally, we conduct a survey across five countries, as well as a more in-depth survey at the authors' institution, to discern students' and educators' perceptions of ChatGPT's use. We find that ChatGPT's performance is comparable, if not superior, to that of students in many courses. Moreover, current AI-text classifiers cannot reliably detect ChatGPT's use in school work, due to their propensity to classify human-written answers as AI-generated, as well as the ease with which AI-generated text can be edited to evade detection. Finally, we find an emerging consensus among students to use the tool, and among educators to treat this as plagiarism. Our findings offer insights that could guide policy discussions addressing the integration of AI into educational frameworks.
研究动机与目标
- 评估ChatGPT在多样化学术学科中相对于大学生的表现。
- 评估使用两种AI文本分类器(GPTZero和OpenAI分类器)检测ChatGPT生成文本的可行性。
- 研究这些分类器在使用Quillbot等工具进行混淆攻击时的脆弱性。
- 考察学生和教育工作者对在学术作业中使用ChatGPT的看法。
- 为学术诚信政策及学生评估框架中的AI使用提供政策制定依据。
提出的方法
- 从纽约大学阿布扎比分校及其他机构收集了32门大学课程的题目及对应的学生作答。
- 使用ChatGPT生成所有课程题目的回答,以与学生提交的答案进行直接对比。
- 在五个国家及纽约大学阿布扎比分校开展调查,评估学生和教师对AI使用的看法。
- 应用两种AI文本分类器——GPTZero和OpenAI分类器——对学生的文本及ChatGPT生成的文本进行分类。
- 通过在多种模式下以最大同义词替换强度使用Quillbot重写ChatGPT输出,实施混淆攻击。
- 通过构建每个分类器的归一化混淆矩阵,计算假阳性和假阴性率。

实验结果
研究问题
- RQ1ChatGPT在广泛学术课程中的表现与大学生相比如何?
- RQ2当前的AI文本分类器在学术提交物中可靠检测ChatGPT生成文本的程度如何?
- RQ3如Quillbot重写等混淆技术在规避AI检测系统方面的有效性如何?
- RQ4学生和教育工作者对在学术作业中使用ChatGPT的实证与规范性看法是什么?
- RQ5这些发现对高等教育学术诚信政策有何影响?
主要发现
- 在32门大学课程中,ChatGPT在28门课程中的表现优于或等同于学生,尤其在人文学科和社会科学领域表现尤为突出。
- GPTZero分类器将32.5%的人工撰写的学生作答错误识别为AI生成,表明其存在较高的假阳性率。
- OpenAI分类器将28.3%的人工作答错误识别为AI生成,进一步凸显检测系统的不可靠性。
- GPTZero将12.5%的ChatGPT回答错误识别为人工撰写,OpenAI分类器则错误识别了15.7%,表明存在显著的假阴性率。
- 使用Quillbot实施的混淆攻击在两种分类器下成功规避检测的比例达68.8%,证明了其高度的规避成功率。
- 调查中,87%的教育工作者认为AI生成内容属于抄袭,而73%的学生表示曾使用ChatGPT完成学术任务。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。