Skip to main content
QUICK REVIEW

[Paper Review] An Empirical Study of AI Generated Text Detection Tools

Arslan Akram|arXiv (Cornell University)|Sep 27, 2023
Artificial Intelligence in Healthcare and EducationMedicine24 references19 citations
TL;DR

The paper evaluates six AI-generated-text detectors across a multi-domain dataset to assess their effectiveness on ChatGPT-produced material.

ABSTRACT

Since ChatGPT has emerged as a major AIGC model, providing high-quality responses across a wide range of applications (including software development and maintenance), it has attracted much interest from many individuals. ChatGPT has great promise, but there are serious problems that might arise from its misuse, especially in the realms of education and public safety. Several AIGC detectors are available, and they have all been tested on genuine text. However, more study is needed to see how effective they are for multi-domain ChatGPT material. This study aims to fill this need by creating a multi-domain dataset for testing the state-of-the-art APIs and tools for detecting artificially generated information used by universities and other research institutions. A large dataset consisting of articles, abstracts, stories, news, and product reviews was created for this study. The second step is to use the newly created dataset to put six tools through their paces. Six different artificial intelligence (AI) text identification systems, including "GPTkit," "GPTZero," "Originality," "Sapling," "Writer," and "Zylalab," have accuracy rates between 55.29 and 97.0%. Although all the tools fared well in the evaluations, originality was particularly effective across the board.

Motivation & Objective

  • Motivate the need to assess detectors on multi-domain ChatGPT material beyond genuine text.
  • Create a large, multi-domain dataset (articles, abstracts, stories, news, product reviews) for detector evaluation.
  • Evaluate state-of-the-art detection tools on the new dataset to measure cross-domain performance.

Proposed method

  • Assemble a large, multi-domain dataset consisting of articles, abstracts, stories, news, and product reviews.
  • Test six AI text detection tools: GPTkit, GPTZero, Originality, Sapling, Writer, and Zylalab.
  • Measure accuracy of each tool on the dataset, reporting a range of performance across tools.

Experimental results

Research questions

  • RQ1How effective are current AI-generated text detectors when evaluated on diverse, non-genuine content from multiple domains?
  • RQ2What is the cross-domain detection accuracy and variability among leading detector tools?
  • RQ3Which detectors show the strongest overall performance and consistency across domains?

Key findings

  • Detector accuracy ranges from 55.29% to 97.0% across tools.
  • Originality performs particularly well across the evaluated tools.
  • Detectors exhibit variability in performance depending on domain and tool.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.