Skip to main content
QUICK REVIEW

[Paper Review] GPT detectors are biased against non-native English writers

Weixin Liang, Mert Yüksekgönül|arXiv (Cornell University)|Apr 6, 2023
Artificial Intelligence in Healthcare and Education37 citations
TL;DR

GPT detectors misclassify many non-native English essays as AI-generated, while native English writings are identified correctly; simple prompts can bypass detectors, raising ethical concerns for education and evaluation.

ABSTRACT

The rapid adoption of generative language models has brought about substantial advancements in digital communication, while simultaneously raising concerns regarding the potential misuse of AI-generated content. Although numerous detection methods have been proposed to differentiate between AI and human-generated content, the fairness and robustness of these detectors remain underexplored. In this study, we evaluate the performance of several widely-used GPT detectors using writing samples from native and non-native English writers. Our findings reveal that these detectors consistently misclassify non-native English writing samples as AI-generated, whereas native writing samples are accurately identified. Furthermore, we demonstrate that simple prompting strategies can not only mitigate this bias but also effectively bypass GPT detectors, suggesting that GPT detectors may unintentionally penalize writers with constrained linguistic expressions. Our results call for a broader conversation about the ethical implications of deploying ChatGPT content detectors and caution against their use in evaluative or educational settings, particularly when they may inadvertently penalize or exclude non-native English speakers from the global discourse. The published version of this study can be accessed at: www.cell.com/patterns/fulltext/S2666-3899(23)00130-7

Motivation & Objective

  • Assess the fairness and robustness of publicly available GPT detectors on native vs. non-native English writing samples.
  • Quantify false positives for non-native writers and false negatives for native writers across detectors.
  • Investigate whether linguistic enhancements or prompts influence detector performance.
  • Examine whether detectors' reliance on perplexity contributes to biases against non-native authors.
  • Provide recommendations for safer, more equitable use of AI content detectors.

Proposed method

  • Evaluate seven off-the-shelf GPT detectors on TOEFL essays (non-native writers) and US 8th-grade essays (native writers).
  • Compute false positive rates and unanimity of AI-generated classifications across detectors.
  • Analyze perplexity differences between groups and correlate with detection outcomes.
  • Use ChatGPT prompts to enhance or simplify language and assess impact on misclassification and perplexity.
  • Test second-round self-edit prompts to assess detector bypass potential.
  • Supplement analysis with a cross-domain check using ICLR 2023 accepted papers to assess perplexity differences by native vs. non-native authors.

Experimental results

Research questions

  • RQ1Do GPT detectors exhibit higher false positive rates for non-native English writing compared to native writing across multiple detectors?
  • RQ2Can linguistic enhancements or prompting strategies mitigate detector bias or alternatively enable bypassing detectors?
  • RQ3Is perplexity a reliable standalone signal for detecting AI-generated text across native/non-native writing?
  • RQ4How do detector biases manifest when applied to academic writing contexts (e.g., conference abstracts) beyond TOEFL/college essays?

Key findings

  • Detectors misclassify over half of non-native TOEFL essays as AI-generated (average false positive rate: 61.22%).
  • Detectors unanimously identified 18 of the 91 TOEFL essays as AI-generated, while 89 of 91 were flagged by at least one detector.
  • Enhancing non-native essays with native-speaker-like word choices via ChatGPT reduced misclassification from 61.22% to 11.77% (1/91 unanimously AI-written).
  • Conversely, simplifying native college essays to resemble non-native writing increased misclassification to 56.65%.
  • Second-round self-edit prompts can drastically reduce detection rates (up to 13% from as high as 100% in some cases) and increase perplexity, demonstrating vulnerability to prompt design.
  • Analyses with ICLR 2023 abstracts show non-native authors have lower perplexity in abstracts, supporting the linkage between linguistic variability and detector bias.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.