Skip to main content
QUICK REVIEW

[论文解读] In which fields can ChatGPT detect journal article quality? An evaluation of REF2021 results

Mike Thelwall, Abdallah Yaghi|arXiv (Cornell University)|Sep 25, 2024
Artificial Intelligence in Healthcare and Education被引用 10
一句话总结

本研究测试 ChatGPT 4o-mini 是否能够在 REF2021 各领域中估算期刊文章质量,通过将其分数与部门平均值进行比较。

ABSTRACT

Time spent by academics on research quality assessment might be reduced if automated approaches can help. Whilst citation-based indicators have been extensively developed and evaluated for this, they have substantial limitations and Large Language Models (LLMs) like ChatGPT provide an alternative approach. This article assesses whether ChatGPT 4o-mini can be used to estimate the quality of journal articles across academia. It samples up to 200 articles from all 34 Units of Assessment (UoAs) in the UK's Research Excellence Framework (REF) 2021, comparing ChatGPT scores with departmental average scores. There was an almost universally positive Spearman correlation between ChatGPT scores and departmental averages, varying between 0.08 (Philosophy) and 0.78 (Psychology, Psychiatry and Neuroscience), except for Clinical Medicine (rho=-0.12). Although other explanations are possible, especially because REF score profiles are public, the results suggest that LLMs can provide reasonable research quality estimates in most areas of science, and particularly the physical and health sciences and engineering, even before citation data is available. Nevertheless, ChatGPT assessments seem to be more positive for most health and physical sciences than for other fields, a concern for multidisciplinary assessments, and the ChatGPT scores are only based on titles and abstracts, so cannot be research evaluations.

研究动机与目标

  • 推动减少学者在研究质量评估上花费的时间。
  • 探索大型语言模型是否能够跨学科估算期刊文章质量。
  • 评估 ChatGPT 派生分数与公认的 REF201가 平均值之间的相关性。

提出的方法

  • 从所有 34 个 REF2021 评估单位(UoAs)抽取最多 200 篇文章。
  • 对每篇文章计算 ChatGPT 分数,并与部门平均分数进行比较。
  • 评估 ChatGPT 分数与部门平均值之间的斯皮尔曼相关性。
  • 分析按领域的变异性并识别显著离群值。
  • 注意,ChatGPT 分数仅基于标题和摘要。

实验结果

研究问题

  • RQ1ChatGPT 派生的分数是否能够在跨学科上近似部门报告的 REF2021 质量分数?
  • RQ2ChatGPT 分数与部门平均值之间的相关性在各领域如何变化?
  • RQ3是否存在 ChatGPT 表现显著较差或较好的领域?
  • RQ4仅使用标题和摘要进行质量评估会带来哪些处于局限性?

主要发现

  • 在大多数领域,ChatGPT 分数与部门平均值之间几乎普遍存在正相关(斯皮尔曼相关)。
  • 相关性范围为 0.08(哲学)到 0.78(心理学、精神病学与神经科学)。
  • 临床医学显示负相关(rho = -0.12)。
  • 对大多数健康与物理科学领域,ChatGPT 估计往往更积极,而对其他领域则较低。
  • ChatGPT 评估仅依赖标题和摘要,限制了其用于完整研究评估的用途。
  • 结果表明,大型语言模型在许多领域,特别是物理与健康科学和工程方面,即使在没有引文数据之前,也能提供合理的质量估计。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。