Skip to main content
QUICK REVIEW

[Paper Review] ChatGPT "contamination": estimating the prevalence of LLMs in the scholarly literature

Andrew Gray|arXiv (Cornell University)|Mar 25, 2024
Artificial Intelligence in Healthcare and Education29 citations
TL;DR

The paper estimates how many scholarly articles in 2023 were assisted by large language models (LLMs) like ChatGPT by using keywords disproportionately present in LLM-generated text, estimating at least 60,000 papers (slightly over 1% of all articles).

ABSTRACT

The use of ChatGPT and similar Large Language Model (LLM) tools in scholarly communication and academic publishing has been widely discussed since they became easily accessible to a general audience in late 2022. This study uses keywords known to be disproportionately present in LLM-generated text to provide an overall estimate for the prevalence of LLM-assisted writing in the scholarly literature. For the publishing year 2023, it is found that several of those keywords show a distinctive and disproportionate increase in their prevalence, individually and in combination. It is estimated that at least 60,000 papers (slightly over 1% of all articles) were LLM-assisted, though this number could be extended and refined by analysis of other characteristics of the papers or by identification of further indicative keywords.

Motivation & Objective

  • Motivate understanding of how widely LLM tools are used in scholarly writing.
  • Propose a keyword-based methodology to detect LLM-assisted writing in the literature.
  • Provide an estimate of the prevalence of LLM-assisted papers for the publishing year 2023.

Proposed method

  • Use keywords known to be disproportionately present in LLM-generated text as indicators.
  • Analyze prevalence of these keywords in scholarly texts for 2023 to detect disproportionate increases.
  • Consider both individual keywords and their combinations to identify LLM-assisted writing signatures.
  • Produce an estimate of the total number of LLM-assisted papers based on keyword signals.

Experimental results

Research questions

  • RQ1Can a set of keywords disproportionately present in LLM-generated text reveal LLM-assisted writing in the scholarly corpus?
  • RQ2What is the estimated share of 2023 scholarly articles that are LLM-assisted based on keyword prevalence?
  • RQ3How do individual keywords and keyword combinations contribute to the detection of LLM contamination?
  • RQ4What refinements would improve precision in estimating LLM-assisted prevalence in literature?

Key findings

  • Several keywords show a distinctive and disproportionate increase in 2023, individually and in combination.
  • Estimated that at least 60,000 papers were LLM-assisted in 2023, equivalent to slightly over 1% of all articles.
  • The estimate could be extended or refined by analyzing additional paper characteristics or identifying more indicative keywords.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.