Skip to main content
QUICK REVIEW

[Paper Review] Can large language models replace humans in the systematic review process? Evaluating GPT-4's efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages

Qusai Khraisha, S. Van Put|arXiv (Cornell University)|Oct 26, 2023
Artificial Intelligence in Healthcare and EducationMedicine37 references11 citations
TL;DR

The study pre-registers and tests GPT-4's autonomous performance in title/abstract screening, full-text screening, and data extraction across peer‑reviewed, grey, and non-English literature, finding that GPT-4 often underperforms humans when accounting for chance and dataset imbalance, but can achieve near-parallel results under highly reliable prompts, especially in full-text screening.

ABSTRACT

Systematic reviews are vital for guiding practice, research, and policy, yet they are often slow and labour-intensive. Large language models (LLMs) could offer a way to speed up and automate systematic reviews, but their performance in such tasks has not been comprehensively evaluated against humans, and no study has tested GPT-4, the biggest LLM so far. This pre-registered study evaluates GPT-4's capability in title/abstract screening, full-text review, and data extraction across various literature types and languages using a 'human-out-of-the-loop' approach. Although GPT-4 had accuracy on par with human performance in most tasks, results were skewed by chance agreement and dataset imbalance. After adjusting for these, there was a moderate level of performance for data extraction, and - barring studies that used highly reliable prompts - screening performance levelled at none to moderate for different stages and languages. When screening full-text literature using highly reliable prompts, GPT-4's performance was 'almost perfect.' Penalising GPT-4 for missing key studies using highly reliable prompts improved its performance even more. Our findings indicate that, currently, substantial caution should be used if LLMs are being used to conduct systematic reviews, but suggest that, for certain systematic review tasks delivered under reliable prompts, LLMs can rival human performance.

Motivation & Objective

  • Assess GPT-4's autonomous performance in title/abstract screening, full-text screening, and data extraction for a systematic review topic.
  • Evaluate GPT-4 on peer‑reviewed, grey, and non-English literature, including grey literature and multilingual sources.
  • Pre-register and document prompt engineering and analysis to understand reliability and bias in LLM-assisted screening.

Proposed method

  • Using GPT-4 via the ChatGPT interface (May–Sept 2023) to screen 300 titles/abstracts and 150 full-texts and to extract data from 30 documents.
  • Testing four inclusion/exclusion prompts for title/abstract screening; adjusting prompts to manage data volume and context; assessing test–retest reliability with 10 studies per criterion.
  • Measuring performance with true positives, true negatives, false positives, and false negatives; reporting sensitivity, specificity, and accuracy.
  • Adjusting for chance agreement and dataset imbalance using Cohen's kappa, PABAK, and weighted kappa to gauge agreement quality.
  • Balancing datasets and reporting literature-type and language-specific performance, including a high-reliability prompt subgroup and non-English/grey literature.
  • Reporting and interpreting an inter-rater reliability benchmark between human reviewers (Cohen’s kappa ~0.77) for context.

Experimental results

Research questions

  • RQ1Can GPT-4 autonomously screen titles/abstracts and full texts with accuracy comparable to human reviewers across different literature types and languages?
  • RQ2How does GPT-4 perform in data extraction across peer‑reviewed, grey, and non-English studies?
  • RQ3What is the impact of prompt reliability and prompt design on GPT-4's screening and extraction performance?
  • RQ4To what extent do chance agreement and dataset balance influence GPT-4's measured performance in systematic reviews?

Key findings

  • GPT-4 showed high reliability for some tasks (e.g., empirical data and refugees) but lower reliability for other concepts like parenting behaviour and protracted refugee situations.
  • Across stages and languages, GPT-4’s sensitivity and specificity varied, with very high specificity (>0.8) generally, and sensitivity ranging from 0.36 to 0.75 depending on literature type and stage.
  • In English peer‑reviewed full-text screening, accuracy was relatively low (0.69) compared with non-English datasets (0.96 for full text) and extraction (0.84 for English peer‑reviewed).
  • A sub-sample with high‑reliability prompts yielded near‑perfect agreement (kappa ~0.85–0.97 when weighted), suggesting prompt quality critically drives performance.
  • Overall, when accounting for imbalance and chance agreement, GPT-4’s performance often lagged behind humans, except under highly reliable prompt conditions in full-text screening where near‑perfect performance was observed.
  • The study emphasizes caution in applying LLMs to systematic reviews broadly, while noting potential for task-specific, prompt-reliable contexts to match human performance.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.