[논문 리뷰] In which fields can ChatGPT detect journal article quality? An evaluation of REF2021 results
본 연구는 ChatGPT 4o-mini가 부서 평균과의 비교를 통해 REF2021 분야 전반에서 저널 기사 품질을 추정할 수 있는지 확인한다.
Time spent by academics on research quality assessment might be reduced if automated approaches can help. Whilst citation-based indicators have been extensively developed and evaluated for this, they have substantial limitations and Large Language Models (LLMs) like ChatGPT provide an alternative approach. This article assesses whether ChatGPT 4o-mini can be used to estimate the quality of journal articles across academia. It samples up to 200 articles from all 34 Units of Assessment (UoAs) in the UK's Research Excellence Framework (REF) 2021, comparing ChatGPT scores with departmental average scores. There was an almost universally positive Spearman correlation between ChatGPT scores and departmental averages, varying between 0.08 (Philosophy) and 0.78 (Psychology, Psychiatry and Neuroscience), except for Clinical Medicine (rho=-0.12). Although other explanations are possible, especially because REF score profiles are public, the results suggest that LLMs can provide reasonable research quality estimates in most areas of science, and particularly the physical and health sciences and engineering, even before citation data is available. Nevertheless, ChatGPT assessments seem to be more positive for most health and physical sciences than for other fields, a concern for multidisciplinary assessments, and the ChatGPT scores are only based on titles and abstracts, so cannot be research evaluations.
연구 동기 및 목표
- 학계의 연구 품질 평가에 소요되는 시간을 줄이는 것을 동기로 삼는다.
- 대형 언어 모델이 여러 학문 분야에 걸쳐 저널 기사 품질을 추정할 수 있는지 탐색한다.
- ChatGPT 파생 점수와 확립된 REF201가 평균 간의 상관 관계를 평가한다.
제안 방법
- 모든 34개 REF2021 평가 부문(UoA)에서 최대 200편의 기사 샘플링.
- 각 기사에 대한 ChatGPT 점수를 계산하고 이를 부서 평균 점수와 비교한다.
- ChatGPT 점수와 부서 평균 간의 Spearman 상관과의 연관성을 평가한다.
- 분야별 변Variation 분석 및 주목할 만한 특이치를 식별한다.
- 참고로 ChatGPT 점수는 제목 및 초록만을 기반으로 한다.
실험 결과
연구 질문
- RQ1ChatGPT 유도 점수가 학과가 보고한 REF2021 품질 점수를 분야 전반에 걸쳐 근사할 수 있는가?
- RQ2ChatGPT 점수와 부서 평균 간의 상관 관계가 분야별로 어떻게 달라지는가?
- RQ3ChatGPT가 특히 성능이 낮거나 높은 분야가 있는가?
- RQ4품질 평가를 제목과 초록만으로 제한할 때의 한계는 무엇인가?
주요 결과
- 대부분의 분야에서 ChatGPT 점수와 부서 평균 간에 거의 전반적으로 양의 Spearman 상관이 있다.
- 상관 관계는 0.08(Philosophy)에서 0.78(Psychology, Psychiatry and Neuroscience)까지이다.
- Clinical Medicine은 음의 상관관계를 보인다(rho = -0.12).
- 대부분의 보건 및 물리과학 분야에서 다른 분야보다 ChatGPT 추정값이 더 긍정적인 경향을 보인다.
- ChatGPT 평가은 제목과 초록에만 의존하므로 전체 연구 평가에 대한 활용이 제한된다.
- 결과는 LLM이 인용 데이터 이전에도 물리 및 보건 과학, 공학의 영역에서 합리적인 품질 추정치를 제공할 수 있음을 시사한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.