Skip to main content
QUICK REVIEW

[论文解读] Automating Systematic Literature Reviews with Natural Language Processing and Text Mining: a Systematic Literature Review

Girish Sundaram, Daniel Berleant|arXiv (Cornell University)|Nov 20, 2022
Software Engineering Research被引用 6
一句话总结

本文通过系统文献综述分析了29项关于使用自然语言处理(NLP)和文本挖掘技术自动化系统文献综述(SLRs)的研究。研究识别出关键的自动化目标——研究选择、数据提取、质量评估和综合分析——同时指出了主流机器学习技术、持续存在的挑战以及当前方法中的空白,特别是在质量评估和综合分析方面,呼吁进一步研究以提升SLR的效率和可扩展性。

ABSTRACT

Objectives: An SLR is presented focusing on text mining based automation of SLR creation. The present review identifies the objectives of the automation studies and the aspects of those steps that were automated. In so doing, the various ML techniques used, challenges, limitations and scope of further research are explained. Methods: Accessible published literature studies that primarily focus on automation of study selection, study quality assessment, data extraction and data synthesis portions of SLR. Twenty-nine studies were analyzed. Results: This review identifies the objectives of the automation studies, steps within the study selection, study quality assessment, data extraction and data synthesis portions that were automated, the various ML techniques used, challenges, limitations and scope of further research. Discussion: We describe uses of NLP/TM techniques to support increased automation of systematic literature reviews. This area has attracted increase attention in the last decade due to significant gaps in the applicability of TM to automate steps in the SLR process. There are significant gaps in the application of TM and related automation techniques in the areas of data extraction, monitoring, quality assessment and data synthesis. There is thus a need for continued progress in this area, and this is expected to ultimately significantly facilitate the construction of systematic literature reviews.

研究动机与目标

  • 调查使用自然语言处理(NLP)和文本挖掘技术在系统文献综述(SLRs)中实现自动化的程度和目标。
  • 识别现有研究中系统文献综述(SLR)流程的哪些步骤——如研究选择、数据提取、质量评估和数据综合——已被实现自动化。
  • 分析自动化研究中采用的机器学习技术,并评估其有效性和局限性。
  • 突出当前自动化中的空白,特别是质量评估和数据综合方面的不足,并提出未来研究的方向。
  • 为研究人员和实践者提供SLR自动化领域最新进展的全面概述,以指导其采用和改进自动化SLR工作流程。

提出的方法

  • 对29项聚焦于使用自然语言处理(NLP)和文本挖掘技术自动化SLR步骤的研究进行了系统文献综述(SLR)。
  • 根据与核心SLR阶段(即研究选择、质量评估、数据提取和数据综合)自动化相关的相关性筛选研究。
  • 根据自动化目标、所用机器学习技术(如NLP模型、聚类、分类)和实现方法对研究进行分类。
  • 将自动化结果映射到特定的SLR阶段,以评估不同阶段的成熟度和覆盖范围。
  • 通过主题分析对所审查文献中的挑战、局限性和研究空白进行综合分析。
  • 使用结构化框架评估自动化技术在整个SLR生命周期中的范围和影响。

实验结果

研究问题

  • RQ1在系统文献综述流程中,哪些阶段最常使用自然语言处理(NLP)和文本挖掘技术实现自动化?
  • RQ2在SLR任务自动化中,主要使用哪些机器学习和自然语言处理(NLP)技术?
  • RQ3现有SLR自动化研究中报告的主要挑战和局限性是什么?
  • RQ4当前SLR中数据提取、质量评估和数据综合的自动化程度如何?
  • RQ5在推进系统文献综述自动化方面,关键的研究空白和未来方向是什么?

主要发现

  • 研究选择是自动化程度最高的阶段,使用NLP模型在去重检测和相关性筛选方面报告了较高的准确性。
  • 数据提取自动化处于中等先进水平,但受限于结构化模式设计和领域特定的标注需求。
  • 质量评估仍是主要挑战,仅有少数研究实现了自动化工具,表明存在显著的研究空白。
  • 数据综合自动化发展不足,NLP在总结或整合跨研究发现方面的应用极少。
  • 大多数研究依赖监督式机器学习技术,如基于BERT的模型和传统分类器,但很少报告交叉验证或基准测试。
  • 尽管已取得进展,SLR的整体自动化仍不完整,其在可扩展性、可解释性和跨领域泛化方面仍存在持续限制。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。