Skip to main content
QUICK REVIEW

[Paper Review] In which fields can ChatGPT detect journal article quality? An evaluation of REF2021 results

Mike Thelwall, Abdallah Yaghi|arXiv (Cornell University)|Sep 25, 2024
Artificial Intelligence in Healthcare and Education10 citations
TL;DR

This study tests whether ChatGPT 4o-mini can estimate journal article quality across REF2021 fields by comparing its scores to departmental averages.

ABSTRACT

Time spent by academics on research quality assessment might be reduced if automated approaches can help. Whilst citation-based indicators have been extensively developed and evaluated for this, they have substantial limitations and Large Language Models (LLMs) like ChatGPT provide an alternative approach. This article assesses whether ChatGPT 4o-mini can be used to estimate the quality of journal articles across academia. It samples up to 200 articles from all 34 Units of Assessment (UoAs) in the UK's Research Excellence Framework (REF) 2021, comparing ChatGPT scores with departmental average scores. There was an almost universally positive Spearman correlation between ChatGPT scores and departmental averages, varying between 0.08 (Philosophy) and 0.78 (Psychology, Psychiatry and Neuroscience), except for Clinical Medicine (rho=-0.12). Although other explanations are possible, especially because REF score profiles are public, the results suggest that LLMs can provide reasonable research quality estimates in most areas of science, and particularly the physical and health sciences and engineering, even before citation data is available. Nevertheless, ChatGPT assessments seem to be more positive for most health and physical sciences than for other fields, a concern for multidisciplinary assessments, and the ChatGPT scores are only based on titles and abstracts, so cannot be research evaluations.

Motivation & Objective

  • Motivate reducing academics' time spent on research quality assessment.
  • Explore whether a large language model can estimate journal article quality across disciplines.
  • Assess the correlation between ChatGPT-derived scores and established REF201가 averages.

Proposed method

  • Sample up to 200 articles from all 34 REF2021 Units of Assessment (UoAs).
  • Compute ChatGPT scores for each article and compare to departmental average scores.
  • Assess association with Spearman correlation between ChatGPT scores and department averages.
  • Analyze field-by-field variation and identify notable outliers.
  • Note that ChatGPT scores are based on titles and abstracts only.

Experimental results

Research questions

  • RQ1Can ChatGPT-derived scores approximate department-reported REF2021 quality scores across disciplines?
  • RQ2How does the correlation between ChatGPT scores and departmental averages vary by field?
  • RQ3Are there fields where ChatGPT performs notably poorly or positively?
  • RQ4What limitations arise from using only titles and abstracts for quality assessment?

Key findings

  • There is an almost universally positive Spearman correlation between ChatGPT scores and departmental averages across most fields.
  • Correlation ranges from 0.08 (Philosophy) to 0.78 (Psychology, Psychiatry and Neuroscience).
  • Clinical Medicine shows a negative correlation (rho = -0.12).
  • ChatGPT estimates tend to be more positive for most health and physical sciences than for other fields.
  • ChatGPT assessments rely solely on titles and abstracts, limiting their use for full research evaluations.
  • Results suggest LLMs can provide reasonable quality estimates in many areas, especially physical and health sciences and engineering, even before citation data.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.