Skip to main content
QUICK REVIEW

[论文解读] Delving into LLM-assisted writing in biomedical publications through excess vocabulary

Dmitry Kobak, Rita González Márquez|arXiv (Cornell University)|Jun 11, 2024
Artificial Intelligence in Healthcare and Education被引用 34
一句话总结

本文提出一种公正的、数据驱动的方法,利用异常词汇量来量化生物医学摘要中的 LLM 辅助写作,估计至少 10%(在某些子语料库中甚至更高)的 2024 年 PubMed 摘要是经过 ChatGPT 等 LLM 处理的。

ABSTRACT

Large language models (LLMs) like ChatGPT can generate and revise text with human-level performance. These models come with clear limitations: they can produce inaccurate information, reinforce existing biases, and be easily misused. Yet, many scientists use them for their scholarly writing. But how wide-spread is such LLM usage in the academic literature? To answer this question for the field of biomedical research, we present an unbiased, large-scale approach: we study vocabulary changes in over 15 million biomedical abstracts from 2010--2024 indexed by PubMed, and show how the appearance of LLMs led to an abrupt increase in the frequency of certain style words. This excess word analysis suggests that at least 13.5% of 2024 abstracts were processed with LLMs. This lower bound differed across disciplines, countries, and journals, reaching 40% for some subcorpora. We show that LLMs have had an unprecedented impact on scientific writing in biomedical research, surpassing the effect of major world events such as the Covid pandemic.

研究动机与目标

  • 在没有真实提示或检测器的情况下衡量 LLM 对科学写作的影响。
  • 识别自 2010–2024 年 1440 万篇 PubMed 摘要中的异常词汇使用模式。
  • 量化 2024 年 ChatGPT 类写作工具如何改变写作风格与词汇。

提出的方法

  • 从 PubMed 摘要构建一个 14.4M × 2.4M 的词出现矩阵。
  • 使用观察到的 2024 年频率和对照假设 2021–22 外推(p, q, r, delta)来定义 excess words。
  • 将 829 个 excess words 注释为 content 或 style,并对词性进行分类。
  • 按领域、国家和期刊分析子组差异。
  • 从词组的频率差异中计算 LLM 使用的下界。
Figure 1: Frequencies of PubMed abstracts containing certain words. Black lines show counterfactual extrapolations from 2021–22 to 2023–24. The first six words are affected by ChatGPT; the last three relate to major events that influenced scientific writing and are shown for comparison.
Figure 1: Frequencies of PubMed abstracts containing certain words. Black lines show counterfactual extrapolations from 2021–22 to 2023–24. The first six words are affected by ChatGPT; the last three relate to major events that influenced scientific writing and are shown for comparison.

实验结果

研究问题

  • RQ1科学摘要中的 excess word usage 能否在没有真实标签的情况下揭示 LLM 介入写作?
  • RQ22024 年的 excess vocabulary 脚印在学科、国家和期刊中有多大?
  • RQ3风格词相较于内容词在 LLM 影响写作中的模式是否不同?
  • RQ4LLM 赋能的写作与历史变迁(如 Covid-19 词汇激增)相比如何?

主要发现

  • 2024 年出现 excess words,且风格词(动词和形容词)显著增加,与 Covid 时代的内容词不同。
  • 研究估计至少 10% 的 2024 年摘要经过 LLM 处理,在某些子语料库中的下界高达 30% 。
  • 两组词(全部 excess words 和一个十个词的互不重叠集合)给出类似的下界,LLM 使用约为 11–12%。
  • 领域与国家层面的异质性明显,计算相关领域及部分非英语国家显示更高的下界。
  • 高度检测的期刊与出版商(如 MDPI、Frontiers)显示更高的 excess 使用,而 Nature/Science/Cell 显示较低的下界。
  • 分析将 LLM 影响的写作在质量和数量上,与以往的词汇变动相比,确认为前所未有。
Figure 2: Words showing increased frequency in 2024. (a) Frequencies in 2024 and frequency ratios ( $r$ ). Both axes are on log-scale. Only a subset of points are labeled for visual clarity. The dashed line shows the threshold defining excess words (see text). Words with $r>90$ are shown at $r=90$ .
Figure 2: Words showing increased frequency in 2024. (a) Frequencies in 2024 and frequency ratios ( $r$ ). Both axes are on log-scale. Only a subset of points are labeled for visual clarity. The dashed line shows the threshold defining excess words (see text). Words with $r>90$ are shown at $r=90$ .

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。