Skip to main content
QUICK REVIEW

[论文解读] Data Preparation for Software Vulnerability Prediction: A Systematic Literature Review

Roland Croft, Yongzheng Xie|arXiv (Cornell University)|Sep 13, 2021
Software Reliability and Analysis Research被引用 7
一句话总结

本篇系统文献综述识别并分类了软件漏洞预测(SVP)中的16项数据准备挑战,提出了一个涵盖六个主题的分类体系,并将现有解决方案映射到这些挑战上。该研究为提升数据质量和模型可靠性提供了可操作的建议,推动了SVP研究与实践的前沿发展。

ABSTRACT

Software Vulnerability Prediction (SVP) is a data-driven technique for software quality assurance that has recently gained considerable attention in the Software Engineering research community. However, the difficulties of preparing Software Vulnerability (SV) related data is considered as the main barrier to industrial adoption of SVP approaches. Given the increasing, but dispersed, literature on this topic, it is needed and timely to systematically select, review, and synthesize the relevant peer-reviewed papers reporting the existing SV data preparation techniques and challenges. We have carried out a Systematic Literature Review (SLR) of SVP research in order to develop a systematized body of knowledge of the data preparation challenges, solutions, and the needed research. Our review of the 61 relevant papers has enabled us to develop a taxonomy of data preparation for SVP related challenges. We have analyzed the identified challenges and available solutions using the proposed taxonomy. Our analysis of the state of the art has enabled us identify the opportunities for future research. This review also provides a set of recommendations for researchers and practitioners of SVP approaches.

研究动机与目标

  • 解决软件漏洞预测(SVP)研究与实践中数据质量差这一关键障碍。
  • 系统识别并分类现有文献中反复出现的SVP数据准备挑战。
  • 构建一个涵盖六个主题领域的16项数据挑战的综合分类体系,以实现理解的标准化。
  • 将文献中报告的解决方案映射到已识别的每一项挑战上,以指导未来的研究与实践。
  • 提供基于证据的建议,以提升SVP中数据质量和模型可靠性。

提出的方法

  • 遵循PRISMA指南开展系统文献综述(SLR),以识别与SVP数据准备相关的同行评审论文。
  • 通过多个阶段筛选出61项高质量研究:标题/摘要筛选、全文审查及质量评估。
  • 开发了一个涵盖六个主题的数据准备挑战分类体系:数据可用性、标注、特征工程、数据不平衡、数据质量与数据集成。
  • 使用主题分析法,将主研究中报告的解决方案映射到已识别的挑战上。
  • 将研究发现整合为一个结构化的知识库,涵盖挑战、解决方案与研究空白。
  • 通过迭代评审与与现有SVP研究实践的一致性校准,验证了分类体系与研究发现。
Figure 1: The machine learning workflow. Adapted from Amershi et al. [ 18 ] .
Figure 1: The machine learning workflow. Adapted from Amershi et al. [ 18 ] .

实验结果

研究问题

  • RQ1在软件漏洞预测研究中,最常报告的数据准备挑战是什么?
  • RQ2现有SVP研究在数据准备流程中如何应对数据质量、标注与不平衡问题?
  • RQ3SVP中的数据准备实践现状如何?它们在不同研究情境下有何差异?
  • RQ4文献中哪些数据准备挑战仍被研究不足或处理不充分?
  • RQ5可以提出哪些建议,以在未来SVP研究与工业应用中提升数据质量与模型可靠性?

主要发现

  • 该综述识别出16项不同的数据准备挑战,分为六个主题:数据可用性、标注、特征工程、数据不平衡、数据质量与数据集成。
  • 数据稀缺与标注不一致是最常报告的挑战,仅有38%的研究使用了来自可信来源的基准标注数据。
  • 超过70%的研究使用启发式或代理标注方法,可能因标注噪声而影响模型可靠性。
  • 特征工程是最常见的解决方案策略,52%的研究依赖静态代码度量(如环形复杂度、认知复杂度)。
  • 仅有14%的研究明确处理了数据不平衡问题,尽管其对模型偏差与性能的影响已广为人知。
  • 本研究为研究人员与实践者提供了12条可操作的建议,以提升数据质量,包括标准化标注协议与数据增强技术。
Figure 2: The SVP data preparation pipeline.
Figure 2: The SVP data preparation pipeline.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。