Skip to main content
QUICK REVIEW

[论文解读] On the Use of a Large Language Model to Support the Conduction of a Systematic Mapping Study: A Brief Report from a Practitioner's View

Carolina Barros, Author|arXiv (Cornell University)|Feb 9, 2026
Artificial Intelligence in Healthcare and Education被引用 0
一句话总结

论文报告了使用大型语言模型来辅助系统映射研究的端到端体验,详述时间节省、准确性、提示调整以及人类监督的必要性。

ABSTRACT

The use of Large Language Models (LLMs) has drawn growing interest within the scientific community. LLMs can handle large volumes of textual data and support methods for evidence synthesis. Although recent studies highlight the potential of LLMs to accelerate screening and data extraction steps in systematic reviews, detailed reports of their practical application throughout the entire process remain scarce. This paper presents an experience report on the conduction of a systematic mapping study with the support of LLMs, describing the steps followed, the necessary adjustments, and the main challenges faced. Positive aspects are discussed, such as (i) the significant reduction of time in repetitive tasks and (ii) greater standardization in data extraction, as well as negative aspects, including (i) considerable effort to build reliable well-structured prompts, especially for less experienced users, since achieving effective prompts may require several iterations and testing, which can partially offset the expected time savings, (ii) the occurrence of hallucinations, and (iii) the need for constant manual verification. As a contribution, this work offers lessons learned and practical recommendations for researchers interested in adopting LLMs in systematic mappings and reviews, highlighting both efficiency gains and methodological risks and limitations to be considered.

研究动机与目标

  • 展示在软件工程中对 SMS 的端到端支持使用 LLMs。
  • 评估在筛选与数据提取中的时间效率与准确性:LLM 辅助 vs 手动方法。
  • 识别将 LLMs 整合到 SMS 工作流时的挑战、风险与调整需求。
  • 为研究者在系统映射与综述中使用 LLMs 提供实用建议与经验教训。

提出的方法

  • 定义与 Kitchenham 与 Charters 与 Wohlin 等指南对齐的协议。
  • 先手工筛选标题/摘要,然后使用结构化提示对 ChatGPT-4 进行比较。
  • 在预定义模板下进行手动与 LLM 支持条件的数据提取。
  • 采用双重核对的验证策略以降低幻觉与不一致性。
  • 在子集上测试额外模型(Gemini PRO、Manus、Copilot)以探索跨模型性能。

实验结果

研究问题

  • RQ1LLM 辅助筛选在时间和准确性方面相对于手工筛选对 SMS 的表现如何?
  • RQ2LLM 辅助数据提取在时间和准确性方面相对于手工提取对 SMS 的表现如何?
  • RQ3将 LLMs 整合到 SMS 工作流时需要哪些实际调整、风险与验证需求?
  • RQ4替代的 LLMs(Gemini PRO、Manus、Copilot)在筛选与提取任务上的表现如何?

主要发现

  • LLM 辅助筛选将时间从约 23 天降至约 9 小时(减少 98%).
  • LLM 辅助提取将时间从约 7 天降至约 1 小时(减少 99%).
  • 筛选准确性在 LLM 下约为 95%(208/219 正确,11 发生幻觉)。
  • 提取准确性在 LLM 下约为 92%(12/13 正确,1 错误)。
  • LLM 输出需要人工核验以降低幻觉并确保一致性。
  • Gemini PRO 在测试子集上筛选与提取均达 90% 的准确率;Manus 在筛选 98%、提取 40%;Copilot 在两项任务中均为 60%。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。