Skip to main content
QUICK REVIEW

[论文解读] Evaluating AI and Human Authorship Quality in Academic Writing through Physics Essays

Will Yeadon, Elise Agra|arXiv (Cornell University)|Mar 8, 2024
Artificial Intelligence in Healthcare and EducationMedicine被引用 3
一句话总结

本研究通过盲评方式评估了300篇简短的物理学科论文——其中一半由人类撰写,另一半由GPT-4生成——结果未发现质量上存在显著差异(p = 0.107)。尽管人类判断的准确率仅略高于随机猜测,ZeroGPT在识别AI文本方面却达到了98%的准确率。作者提出将50%的AI内容作为分类为人类撰写的实用界限,以在人工智能整合与学术诚信之间取得平衡。

ABSTRACT

This study evaluates $n = 300$ short-form physics essay submissions, equally divided between student work submitted before the introduction of ChatGPT and those generated by OpenAI's GPT-4. In blinded evaluations conducted by five independent markers who were unaware of the origin of the essays, we observed no statistically significant differences in scores between essays authored by humans and those produced by AI (p-value $= 0.107$, $α$ = 0.05). Additionally, when the markers subsequently attempted to identify the authorship of the essays on a 4-point Likert scale - from `Definitely AI' to `Definitely Human' - their performance was only marginally better than random chance. This outcome not only underscores the convergence of AI and human authorship quality but also highlights the difficulty of discerning AI-generated content solely through human judgment. Furthermore, the effectiveness of five commercially available software tools for identifying essay authorship was evaluated. Among these, ZeroGPT was the most accurate, achieving a 98% accuracy rate and a precision score of 1.0 when its classifications were reduced to binary outcomes. This result is a source of potential optimism for maintaining assessment integrity. Finally, we propose that texts with $\leq 50\%$ AI-generated content should be considered the upper limit for classification as human-authored, a boundary inclusive of a future with ubiquitous AI assistance whilst also respecting human-authorship.

研究动机与目标

  • 评估AI生成的论文在物理学科语境下是否与人类撰写的论文质量相当。
  • 评估人类评分者在学术写作中区分AI与人类作者身份的能力。
  • 测试商业AI检测工具在识别学术论文中AI生成内容方面的有效性。
  • 建立人类撰写的学术作品中可接受AI内容的实用阈值。
  • 为高等教育中AI使用提供教育政策建议,同时维护学术诚信。

提出的方法

  • 由五名不知晓作者身份的独立评分者对300篇简短物理论文(150篇人类撰写,150篇GPT-4生成)进行盲评。
  • 论文选自一门物理课程的形成性与总结性评估,主题涵盖物理学史、哲学与伦理学。
  • 评分者依据标准化标准对论文的学术质量进行评分,作者身份在评分后才被分配,以避免偏见。
  • 后续任务要求评分者在四点李克特量表上将每篇论文分类为“肯定为AI生成”、“可能为AI生成”、“可能为人撰写”或“肯定为人撰写”。
  • 评估了五种商业AI检测工具(包括ZeroGPT)在识别AI生成内容方面的准确性。
  • 提出将≤50%的AI生成内容作为分类为人类撰写的实用边界。

实验结果

研究问题

  • RQ1AI生成与人类撰写的物理学科论文在质量上是否存在统计学上的显著差异?
  • RQ2在盲评环境下,人类评分者能否可靠地区分AI与人类撰写的学术论文?
  • RQ3商业可用的AI检测工具在识别AI生成的学术内容方面效果如何?
  • RQ4在不损害作者身份完整性的前提下,人类撰写的学术作品中可接受的AI内容比例是多少?
  • RQ5可建立何种政策指导原则,以在高等教育中实现AI辅助与学术诚信之间的平衡?

主要发现

  • 未发现人类与AI生成论文在质量上存在显著差异(p = 0.107,α = 0.05)。
  • 人类评分者识别AI作者身份的能力仅略高于随机猜测,准确率未显著高于50%。
  • ZeroGPT在二元分类下检测AI生成内容的准确率达到98%,精确度为1.0。
  • 研究发现,AI生成的论文在内容、结构和学术严谨性方面均达到人类标准,涵盖多种哲学与历史主题的物理议题。
  • 商业改写工具无法有效规避检测,表明需通过人工编辑才能绕过AI检测系统。
  • 作者建议将AI生成内容占比≤50%的作品分类为人类撰写,为未来学术工作提供一种实用且包容的标准。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。