Skip to main content
QUICK REVIEW

[Paper Review] AI, write an essay for me: A large-scale comparison of human-written versus ChatGPT-generated essays

Steffen Herbold, Annette Hautli-Janisz|arXiv (Cornell University)|Apr 24, 2023
Artificial Intelligence in Healthcare and Education17 citations
TL;DR

The study systematically compares human-written and ChatGPT-generated argumentative essays and finds that ChatGPT (especially GPT-4) outperforms humans in overall quality, with distinct linguistic patterns across models.

ABSTRACT

Background: Recently, ChatGPT and similar generative AI models have attracted hundreds of millions of users and become part of the public discourse. Many believe that such models will disrupt society and will result in a significant change in the education system and information generation in the future. So far, this belief is based on either colloquial evidence or benchmarks from the owners of the models -- both lack scientific rigour. Objective: Through a large-scale study comparing human-written versus ChatGPT-generated argumentative student essays, we systematically assess the quality of the AI-generated content. Methods: A large corpus of essays was rated using standard criteria by a large number of human experts (teachers). We augment the analysis with a consideration of the linguistic characteristics of the generated essays. Results: Our results demonstrate that ChatGPT generates essays that are rated higher for quality than human-written essays. The writing style of the AI models exhibits linguistic characteristics that are different from those of the human-written essays, e.g., it is characterized by fewer discourse and epistemic markers, but more nominalizations and greater lexical diversity. Conclusions: Our results clearly demonstrate that models like ChatGPT outperform humans in generating argumentative essays. Since the technology is readily available for anyone to use, educators must act immediately. We must re-invent homework and develop teaching concepts that utilize these AI models in the same way as math utilized the calculator: teach the general concepts first and then use AI tools to free up time for other learning objectives.

Motivation & Objective

  • Assess the quality of AI-generated argumentative essays versus human-written ones using a large pool of expert raters (teachers).
  • Characterize linguistic differences between human and AI-generated essays across two ChatGPT versions (GPT-3.5 and GPT-4).
  • Provide statistically rigorous analysis of essay quality with reliability checks and linguistic feature correlations.

Proposed method

  • Collect a large corpus of student essays (human-written) on 90 topics from an online forum.
  • Prompt ChatGPT-3 and ChatGPT-4 with a basic zero-shot prompt to generate ~200-word essays for the same topics.
  • Have 108 teachers rate 658 ratings across 270 essays on seven criteria using a seven-point Likert scale and compute inter-rater reliability.
  • Perform computational linguistic analysis on lexical diversity, syntactic complexity, nominalisation, modals, epistemic markers, and discourse markers.
  • Use Wilcoxon signed-rank tests with Holm-Bonferroni correction for multiple comparisons and report Cohen’s d as effect size; bootstrap-based confidence intervals.
  • Replicate analysis with an available replication package.

Experimental results

Research questions

  • RQ1RQ1: How good is ChatGPT based on GPT-3 and GPT-4 at writing argumentative student essays?
  • RQ2RQ2: How do AI-generated essays compare to essays written by humans?
  • RQ3RQ3: What linguistic devices are characteristic of human versus AI-generated content?

Key findings

  • ChatGPT-generated essays are rated higher in quality than human-written essays across all criteria, with GPT-4 outperforming GPT-3.5.
  • GPT-4 shows higher performance in logical structure, language complexity, vocabulary richness, and text linking compared to GPT-3.5.
  • Humans use more modals and epistemic markers, while GPT models use more nominalizations and show greater sentence complexity.
  • Linguistic diversity improves over time, with GPT-4 displaying higher diversity than humans, whereas GPT-3.5 lags behind humans in diversity.
  • Differences between GPT-4 and GPT-3.5 are significant for logic, vocabulary linking, and complexity, indicating broad improvements in GPT-4.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.