Skip to main content
QUICK REVIEW

[论文解读] Is ChatGPT Transforming Academics' Writing Style?

Mingmeng Geng, Roberto Trotta|arXiv (Cornell University)|Apr 12, 2024
Artificial Intelligence in Healthcare and Education被引用 5
一句话总结

该论文分析了百万条 arXiv 摘要,通过词频变化来检测 ChatGPT 印记的写作风格,按类别和时间估计 ChatGPT 的影响,其中 CS 显示最强效应(在简单提示下约 35%)。

ABSTRACT

Based on one million arXiv papers submitted from May 2018 to January 2024, we assess the textual density of ChatGPT's writing style in their abstracts through a statistical analysis of word frequency changes. Our model is calibrated and validated on a mixture of real abstracts and ChatGPT-modified abstracts (simulated data) after a careful noise analysis. The words used for estimation are not fixed but adaptive, including those with decreasing frequency. We find that large language models (LLMs), represented by ChatGPT, are having an increasing impact on arXiv abstracts, especially in the field of computer science, where the fraction of LLM-style abstracts is estimated to be approximately 35%, if we take the responses of GPT-3.5 to one simple prompt, "revise the following sentences", as a baseline. We conclude with an analysis of both positive and negative aspects of the penetration of LLMs into academics' writing style.

研究动机与目标

  • 激励并量化 ChatGPT 是否影响 arXiv 摘要中的学术写作风格。
  • 开发一个统计框架,以随时间检测类似 ChatGPT 的词频指纹。
  • 使用真实摘要和 ChatGPT 修改(模拟)摘要来标定并验证该方法。
  • 估计在各学科和时间段中受 ChatGPT 影响文本的密度。
  • 讨论 ChatGPT 进入学术写作的影响、收益与风险。

提出的方法

  • 定义一个变化因子 R_i,用以衡量随时间的词频变化(方程 1)。
  • 通过用简单提示润色真实摘要来进行 ChatGPT 驱动的仿真,以估计词变更率 r̂_ij(方程 2)。
  • 用 η_j(t) 表示受影响摘要份额的项来建模 ChatGPT 的影响(方程 5)。
  • 通过 δ_ij 引入噪声,并构建带有偏差感知的损失 L_j,t(η_j) 以估计 η_j(方程 18–23)。
  • 通过在不同提示和混合比下对单词集合 I_j 进行标定并测试鲁棒性(方程 35–37)。
  • 使用前 ChatGPT 时期对 f*_ij(t) 进行标定,并用 GPT-3.5 驱动的仿真进行验证(第 4 节)。
Figure 1: The 12 words with the highest change rate $R_{i}$ and satisfying $\max_{t}(f_{i}(t))>500$ . The vertical red dashed line demarcates the first time period after ChatGPT’s release.
Figure 1: The 12 words with the highest change rate $R_{i}$ and satisfying $\max_{t}(f_{i}(t))>500$ . The vertical red dashed line demarcates the first time period after ChatGPT’s release.

实验结果

研究问题

  • RQ1词频中的统计特征是否能揭示 ChatGPT 对 arXiv 摘要的影响?
  • RQ2ChatGPT 如何在不同学科和随时间的变化中影响单词使用?
  • RQ3在不同领域,尤其是 CS,ChatGPT 风格写作的密度估计是多少?
  • RQ4对不同提示和校准选择,估计结果的鲁棒性如何?
  • RQ5通过词频来衡量 ChatGPT 影响时有哪些局限性和潜在偏差?

主要发现

  • 就释放后在 arXiv 摘要中可以检测到 ChatGPT 风格文本的渗透,计算机科学显示出最强的 uptake。
  • 在 CS 中,使用简单提示基线(“revise the following sentences”)的估计影响约为 35%。
  • 词频变化同时反映主题趋势(如 COVID-19、LLMs、AI)和非主题变化(如功能词“are”/“is”)。
  • 在多个类别(CS、math、astro、cond-mat)中,像“significant”这类词的 simulated ChatGPT 处理显示出显著增长。
  • 基于校准的、透明的频率分析方法可以在不依赖黑箱检测器的情况下量化 ChatGPT 的影响。
Figure 2: Examples of words with rapidly growing frequency in arXiv abstracts.
Figure 2: Examples of words with rapidly growing frequency in arXiv abstracts.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。