Skip to main content
QUICK REVIEW

[Paper Review] Delving into LLM-assisted writing in biomedical publications through excess vocabulary

Dmitry Kobak, Rita González Márquez|arXiv (Cornell University)|Jun 11, 2024
Artificial Intelligence in Healthcare and Education34 citations
TL;DR

The paper introduces an unbiased, data-driven approach using excess word usage to quantify LLM-assisted writing in biomedical abstracts, estimating that at least 10% (and up to higher in some subcorpora) of 2024 PubMed abstracts were processed with LLMs like ChatGPT.

ABSTRACT

Large language models (LLMs) like ChatGPT can generate and revise text with human-level performance. These models come with clear limitations: they can produce inaccurate information, reinforce existing biases, and be easily misused. Yet, many scientists use them for their scholarly writing. But how wide-spread is such LLM usage in the academic literature? To answer this question for the field of biomedical research, we present an unbiased, large-scale approach: we study vocabulary changes in over 15 million biomedical abstracts from 2010--2024 indexed by PubMed, and show how the appearance of LLMs led to an abrupt increase in the frequency of certain style words. This excess word analysis suggests that at least 13.5% of 2024 abstracts were processed with LLMs. This lower bound differed across disciplines, countries, and journals, reaching 40% for some subcorpora. We show that LLMs have had an unprecedented impact on scientific writing in biomedical research, surpassing the effect of major world events such as the Covid pandemic.

Motivation & Objective

  • Measure LLM influence on scientific writing without ground-truth prompts or detectors.
  • Identify excess word usage patterns in 14.4 million PubMed abstracts from 2010–2024.
  • Quantify how ChatGPT-like writing tools changed writing styles and vocabulary in 2024.

Proposed method

  • Construct a 14.4M × 2.4M word occurrence matrix from PubMed abstracts.
  • Define excess words using observed 2024 frequencies and counterfactual 2021–22 extrapolations (p, q, r, delta).
  • Annotate 829 excess words as content or style and classify parts of speech.
  • Analyze subgroup variations by field, country, and journal.
  • Compute lower bounds on LLM usage from frequency gaps across word groups.
Figure 1: Frequencies of PubMed abstracts containing certain words. Black lines show counterfactual extrapolations from 2021–22 to 2023–24. The first six words are affected by ChatGPT; the last three relate to major events that influenced scientific writing and are shown for comparison.
Figure 1: Frequencies of PubMed abstracts containing certain words. Black lines show counterfactual extrapolations from 2021–22 to 2023–24. The first six words are affected by ChatGPT; the last three relate to major events that influenced scientific writing and are shown for comparison.

Experimental results

Research questions

  • RQ1Can excess word usage in scientific abstracts reveal LLM-assisted writing without ground-truth labeling?
  • RQ2How large is the 2024 excess vocabulary footprint across disciplines, countries, and journals?
  • RQ3Do style words show different patterns than content words in LLM-influenced writing?
  • RQ4How does LLM-induced writing compare to historical shifts like the Covid-19 vocabulary surge?

Key findings

  • Excess words emerged in 2024, with style words (verbs and adjectives) increasing markedly, unlike Covid-era content words.
  • The study estimates at least 10% of 2024 abstracts were LLM-processed, with a lower bound up to 30% in some subcorpora.
  • Two word groups (all excess words and a ten-word non-overlapping set) yield similar lower bounds around 11–12% for LLM usage.
  • Field- and country-level heterogeneity is pronounced, with computation and some non-English-speaking countries showing higher bounds.
  • Highly detected journals and publishers (e.g., MDPI, Frontiers) exhibit larger excess usage, while Nature/Science/Cell show lower bounds.
  • The analysis places LLM-influenced writing as unprecedented in both quality and quantity relative to prior vocabulary shifts.
Figure 2: Words showing increased frequency in 2024. (a) Frequencies in 2024 and frequency ratios ( $r$ ). Both axes are on log-scale. Only a subset of points are labeled for visual clarity. The dashed line shows the threshold defining excess words (see text). Words with $r>90$ are shown at $r=90$ .
Figure 2: Words showing increased frequency in 2024. (a) Frequencies in 2024 and frequency ratios ( $r$ ). Both axes are on log-scale. Only a subset of points are labeled for visual clarity. The dashed line shows the threshold defining excess words (see text). Words with $r>90$ are shown at $r=90$ .

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.