Skip to main content
QUICK REVIEW

[Paper Review] Evaluating AI and Human Authorship Quality in Academic Writing through Physics Essays

Will Yeadon, Elise Agra|arXiv (Cornell University)|Mar 8, 2024
Artificial Intelligence in Healthcare and EducationMedicine3 citations
TL;DR

This study evaluates 300 short physics essays—equally split between human and GPT-4-generated—through blinded peer assessment, finding no statistically significant difference in quality (p = 0.107). Despite high human judgment accuracy only slightly above chance, ZeroGPT achieved 98% accuracy in detecting AI text. The authors propose a 50% AI content threshold as a practical boundary for classifying work as human-authored, balancing AI integration with academic integrity.

ABSTRACT

This study evaluates $n = 300$ short-form physics essay submissions, equally divided between student work submitted before the introduction of ChatGPT and those generated by OpenAI's GPT-4. In blinded evaluations conducted by five independent markers who were unaware of the origin of the essays, we observed no statistically significant differences in scores between essays authored by humans and those produced by AI (p-value $= 0.107$, $α$ = 0.05). Additionally, when the markers subsequently attempted to identify the authorship of the essays on a 4-point Likert scale - from `Definitely AI' to `Definitely Human' - their performance was only marginally better than random chance. This outcome not only underscores the convergence of AI and human authorship quality but also highlights the difficulty of discerning AI-generated content solely through human judgment. Furthermore, the effectiveness of five commercially available software tools for identifying essay authorship was evaluated. Among these, ZeroGPT was the most accurate, achieving a 98% accuracy rate and a precision score of 1.0 when its classifications were reduced to binary outcomes. This result is a source of potential optimism for maintaining assessment integrity. Finally, we propose that texts with $\leq 50\%$ AI-generated content should be considered the upper limit for classification as human-authored, a boundary inclusive of a future with ubiquitous AI assistance whilst also respecting human-authorship.

Motivation & Objective

  • To assess whether AI-generated essays match the quality of human-written essays in a physics context.
  • To evaluate the ability of human markers to distinguish between AI and human authorship in academic writing.
  • To test the effectiveness of commercial AI detection tools in identifying AI-generated content in academic essays.
  • To establish a practical threshold for acceptable AI content in human-authored academic work.
  • To inform educational policy on AI use in assessment while preserving academic integrity.

Proposed method

  • Blind evaluation of 300 short-form physics essays (150 human, 150 GPT-4-generated) by five independent markers unaware of authorship.
  • Essays were drawn from a physics course’s formative and summative assessments, with topics in the history, philosophy, and ethics of physics.
  • Markers rated essays on academic quality using standardized criteria, with authorship origin assigned post-evaluation to prevent bias.
  • A follow-up task required markers to classify each essay as 'Definitely AI', 'Probably AI', 'Probably Human', or 'Definitely Human' on a 4-point Likert scale.
  • Five commercial AI detection tools (including ZeroGPT) were evaluated for accuracy in identifying AI-generated content.
  • A threshold of ≤50% AI-generated content was proposed as a practical boundary for classifying work as human-authored.

Experimental results

Research questions

  • RQ1Is there a statistically significant difference in quality between AI-generated and human-written physics essays?
  • RQ2Can human markers reliably distinguish between AI and human-authored academic essays in a blinded setting?
  • RQ3How effective are commercially available AI detection tools in identifying AI-generated academic content?
  • RQ4What proportion of AI-generated content can be considered acceptable in a human-authored academic work without compromising authorship integrity?
  • RQ5What policy guidelines can be established to balance AI assistance with academic integrity in higher education?

Key findings

  • No statistically significant difference in essay quality was found between human and AI-generated submissions (p = 0.107, α = 0.05).
  • Human markers' ability to identify AI authorship was only marginally better than random chance, with accuracy not significantly above 50%.
  • ZeroGPT achieved 98% accuracy and 1.0 precision in detecting AI-generated content when reduced to binary classification.
  • The study found that AI-generated essays matched human standards in content, structure, and academic rigor across a range of philosophical and historical physics topics.
  • Commercial paraphrasing tools were ineffective at evading detection, indicating that human editing is required to bypass AI detection systems.
  • The authors propose that works with ≤50% AI-generated content should be classified as human-authored, offering a practical and inclusive standard for future academic work.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.