[Paper Review] Assessing the societal influence of academic research with ChatGPT: Impact case study evaluations
This study evaluates whether ChatGPT can effectively assess the societal impact of academic research by analyzing 6,220 Impact Case Studies (ICS) from the UK's REF2021 using GPT-4o-mini. It finds that inputting only the title and summary while modifying evaluation guidelines yields high correlations (r = 0.56) with expert scores, especially in fields like Psychology and Sport Sciences, demonstrating ChatGPT’s viability as a support tool for impact assessment.
Academics and departments are sometimes judged by how their research has benefitted society. For example, the UK Research Excellence Framework (REF) assesses Impact Case Studies (ICS), which are five-page evidence-based claims of societal impacts. This study investigates whether ChatGPT can evaluate societal impact claims and therefore potentially support expert human assessors. For this, various parts of 6,220 public ICS from REF2021 were fed to ChatGPT 4o-mini along with the REF2021 evaluation guidelines, comparing the results with published departmental average ICS scores. The results suggest that the optimal strategy for high correlations with expert scores is to input the title and summary of an ICS but not the remaining text, and to modify the original REF guidelines to encourage a stricter evaluation. The scores generated by this approach correlated positively with departmental average scores in all 34 Units of Assessment (UoAs), with values between 0.18 (Economics and Econometrics) and 0.56 (Psychology, Psychiatry and Neuroscience). At the departmental level, the corresponding correlations were higher, reaching 0.71 for Sport and Exercise Sciences, Leisure and Tourism. Thus, ChatGPT-based ICS evaluations are simple and viable to support or cross-check expert judgments, although their value varies substantially between fields.
Motivation & Objective
- To investigate whether large language models like ChatGPT can evaluate the societal impact of academic research as a support tool for expert assessors.
- To determine the optimal input configuration (e.g., title, summary, full text) for maximizing alignment with expert-graded impact scores.
- To assess how modifications to evaluation guidelines influence model performance and correlation with official departmental scores.
- To examine field-specific variations in model performance across 34 Units of Assessment (UoAs) in the UK's Research Excellence Framework (REF).
Proposed method
- Retrieved 6,220 public Impact Case Studies (ICS) from the UK’s REF2021 database.
- Fed only the title and summary of each ICS into GPT-4o-mini, alongside modified versions of the official REF2021 evaluation guidelines.
- Modified the REF2021 guidelines to encourage stricter, more consistent evaluation by the model.
- Compared the model-generated scores against published departmental average ICS scores across all 34 Units of Assessment (UoAs).
- Calculated Pearson correlation coefficients between model scores and expert scores at both the UoA and departmental levels.
- Analyzed field-specific performance variations to identify high- and low-performing disciplines.
Experimental results
Research questions
- RQ1Can ChatGPT accurately evaluate the societal impact of academic research when provided with only the title and summary of an Impact Case Study?
- RQ2How does modifying the REF2021 evaluation guidelines affect the correlation between ChatGPT-generated scores and expert-graded scores?
- RQ3What is the level of agreement between ChatGPT-generated impact scores and official departmental average scores across different academic disciplines?
- RQ4How does model performance vary across the 34 Units of Assessment in the UK’s Research Excellence Framework (REF)?
- RQ5To what extent can ChatGPT serve as a reliable support tool for expert impact assessment in research evaluation?
Key findings
- The highest correlation between ChatGPT-generated scores and expert scores was observed when only the title and summary of each Impact Case Study were input, with a correlation of 0.56 in the Psychology, Psychiatry and Neuroscience unit.
- Correlations ranged from 0.18 in Economics and Econometrics to 0.56 in Psychology, Psychiatry and Neuroscience, indicating significant field-dependent performance variation.
- At the departmental level, the highest correlation reached 0.71 for the Sport and Exercise Sciences, Leisure and Tourism unit, suggesting strong alignment in specific disciplines.
- Modifying the REF2021 guidelines to encourage stricter evaluation improved the model’s consistency and alignment with expert judgments.
- The model demonstrated consistent performance across all 34 Units of Assessment, with positive correlations in every field, indicating broad applicability.
- The results suggest that ChatGPT can serve as a viable, scalable tool for supporting or cross-checking expert impact assessments in research evaluation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.