[Paper Review] On the Use of a Large Language Model to Support the Conduction of a Systematic Mapping Study: A Brief Report from a Practitioner's View
The paper reports an end-to-end experience of using LLMs to assist a systematic mapping study, detailing time savings, accuracy, prompt adjustments, and the need for human oversight.
The use of Large Language Models (LLMs) has drawn growing interest within the scientific community. LLMs can handle large volumes of textual data and support methods for evidence synthesis. Although recent studies highlight the potential of LLMs to accelerate screening and data extraction steps in systematic reviews, detailed reports of their practical application throughout the entire process remain scarce. This paper presents an experience report on the conduction of a systematic mapping study with the support of LLMs, describing the steps followed, the necessary adjustments, and the main challenges faced. Positive aspects are discussed, such as (i) the significant reduction of time in repetitive tasks and (ii) greater standardization in data extraction, as well as negative aspects, including (i) considerable effort to build reliable well-structured prompts, especially for less experienced users, since achieving effective prompts may require several iterations and testing, which can partially offset the expected time savings, (ii) the occurrence of hallucinations, and (iii) the need for constant manual verification. As a contribution, this work offers lessons learned and practical recommendations for researchers interested in adopting LLMs in systematic mappings and reviews, highlighting both efficiency gains and methodological risks and limitations to be considered.
Motivation & Objective
- Demonstrate end-to-end use of LLMs to support an SMS in software engineering.
- Evaluate time efficiency and accuracy of LLM-assisted screening and data extraction versus manual methods.
- Identify challenges, risks, and adjustments required when integrating LLMs into SMS workflows.
- Provide practical recommendations and lessons learned for researchers using LLMs in systematic mappings and reviews.
Proposed method
- Define a protocol aligned with Kitchenham and Charters and Wohlin et al. guidelines.
- Screen titles/abstracts manually first, then with structured prompts to ChatGPT-4 for comparison.
- Perform data extraction under manual and LLM-supported conditions using predefined templates.
- Apply a double-checking verification strategy to mitigate hallucinations and discrepancies.
- Test additional models (Gemini PRO, Manus, Copilot) on subsets to explore cross-model performance.
Experimental results
Research questions
- RQ1How does LLM-assisted screening compare to manual screening in time and accuracy for an SMS?
- RQ2How does LLM-assisted data extraction compare to manual extraction in time and accuracy for an SMS?
- RQ3What are the practical adjustments, risks, and verification needs when integrating LLMs into SMS workflows?
- RQ4How do alternative LLMs (Gemini PRO, Manus, Copilot) perform on screening and extraction tasks?
Key findings
- LLM-assisted screening reduced time from ~23 days to ~9 hours (98% reduction).
- LLM-assisted extraction reduced time from ~7 days to ~1 hour (99% reduction).
- Screening accuracy with LLMs was ~95% (208/219 correct, 11 hallucinations).
- Extraction accuracy with LLMs was ~92% (12/13 correct, 1 error).
- LLM outputs required human verification to mitigate hallucinations and ensure consistency.
- Gemini PRO showed 90% accuracy in both screening and extraction on tested subsets; Manus showed 98% in screening and 40% in extraction; Copilot showed 60% in both tasks.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.