[Paper Review] Assessing the Ability of ChatGPT to Screen Articles for Systematic Reviews
The paper evaluates ChatGPT's consistency, classification performance, and generalizability for screening articles in systematic reviews, comparing it to traditional classifiers, and discusses integration considerations.
By organizing knowledge within a research field, Systematic Reviews (SR) provide valuable leads to steer research. Evidence suggests that SRs have become first-class artifacts in software engineering. However, the tedious manual effort associated with the screening phase of SRs renders these studies a costly and error-prone endeavor. While screening has traditionally been considered not amenable to automation, the advent of generative AI-driven chatbots, backed with large language models is set to disrupt the field. In this report, we propose an approach to leverage these novel technological developments for automating the screening of SRs. We assess the consistency, classification performance, and generalizability of ChatGPT in screening articles for SRs and compare these figures with those of traditional classifiers used in SR automation. Our results indicate that ChatGPT is a viable option to automate the SR processes, but requires careful considerations from developers when integrating ChatGPT into their SR tools.
Motivation & Objective
- Motivate the need to automate the screening phase of systematic reviews in software engineering.
- Assess ChatGPT’s consistency, performance, and generalizability in screening tasks.
- Compare ChatGPT with traditional machine learning baselines used in SR automation.
- Provide guidance on integrating ChatGPT into SR tools based on empirical findings.
Proposed method
- Use ReLiS-grounded SR corpora as ground truth for screening decisions.
- Establish baselines with traditional classifiers (LR, RF, CNB, SVC) trained on article titles/abstracts with Word2Vec features.
- Engineer prompts for ChatGPT and evaluate its screening decisions against ground truth.
- Compare ChatGPT results to baselines using imbalanced-data performance metrics (e.g., MCC, F2, balanced accuracy).
- Assess consistency with Fleiss’ Kappa across multiple runs to measure stability of ChatGPT decisions.
Experimental results
Research questions
- RQ1RQ1: How consistent are ChatGPT’s screening decisions for specific articles across runs?
- RQ2RQ2: How does ChatGPT’s classification performance compare to traditional SR automation classifiers?
- RQ3RQ3: How generalizable are ChatGPT’s screening decisions across different software engineering SR datasets?
Key findings
- ChatGPT can match the performance of traditional machine learning methods in SR screening without additional training.
- LLMs like ChatGPT show potential to revolutionize SR automation, but require careful integration considerations for tool developers.
- Diverse SR datasets with varying inclusion/exclusion ratios and conflicts were used to test generalizability, highlighting the need for robust prompt engineering.
- The study uses ground-truth data from ReLiS projects, ensuring reliable evaluation of inclusion/exclusion decisions.
- Prompts and hyperparameters (temperature, token limits) are critical for consistent, minimal-responsed outcomes (Include/Exclude).
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.