[论文解读] A Systematic Review of Unsupervised Learning Techniques for Software Defect Prediction
本篇系统性综述通过49项研究和2,456项实验结果的元分析,评估了用于软件缺陷预测的无监督学习技术。研究发现,无监督模型,尤其是模糊C-均值(FCM)和模糊SOM(FSOM),其性能与有监督模型相当,但报告质量与实验一致性方面普遍存在严重问题,影响了结果的可靠性。
Background: Unsupervised machine learners have been increasingly applied to software defect prediction. It is an approach that may be valuable for software practitioners because it reduces the need for labeled training data. Objective: Investigate the use and performance of unsupervised learning techniques in software defect prediction. Method: We conducted a systematic literature review that identified 49 studies containing 2456 individual experimental results, which satisfied our inclusion criteria published between January 2000 and March 2018. In order to compare prediction performance across these studies in a consistent way, we (re-)computed the confusion matrices and employed the Matthews Correlation Coefficient (MCC) as our main performance measure. Results: Our meta-analysis shows that unsupervised models are comparable with supervised models for both within-project and cross-project prediction. Among the 14 families of unsupervised model, Fuzzy CMeans (FCM) and Fuzzy SOMs (FSOMs) perform best. In addition, where we were able to check, we found that almost 11% (262/2456) of published results (contained in 16 papers) were internally inconsistent and a further 33% (823/2456) provided insufficient details for us to check. Conclusion: Although many factors impact the performance of a classifier, e.g., dataset characteristics, broadly speaking, unsupervised classifiers do not seem to perform worse than the supervised classifiers in our review. However, we note a worrying prevalence of (i) demonstrably erroneous experimental results, (ii) undemanding benchmarks and (iii) incomplete reporting. We therefore encourage researchers to be comprehensive in their reporting.
研究动机与目标
- 评估无监督学习技术在软件缺陷预测中的性能,与有监督方法进行比较。
- 识别在项目内和跨项目设置下,最有效的无监督模型族。
- 评估无监督缺陷预测研究中实验报告的质量,重点关注完整性和一致性。
- 为从业者和研究人员提供关于无监督缺陷预测模型可行性与可靠性的实用指导。
- 揭示实验设计与报告中的系统性问题,这些问题损害了该领域的可重现性与有效性。
提出的方法
- 使用五个主要数据库(Web of Science、ACM、IEEE Xplore、ScienceDirect、SpringerLink)从2000年1月到2018年3月开展系统性文献综述。
- 应用预设的纳入标准,识别出49项关于无监督缺陷预测的原始研究,共获得2,456项独立实验结果。
- 根据报告的性能指标重新计算混淆矩阵,以确保各研究之间的一致性,并实现公平比较。
- 采用马修斯相关系数(MCC)作为元分析的主要性能度量,以支持跨研究比较。
- 进行文献计量学和质量评估分析,以评估报告的完整性、一致性及研究趋势。
- 将模型分类为14类家族和6种聚类标签方法,以比较不同算法族之间的性能。
实验结果
研究问题
- RQ12000年至2018年间,无监督软件缺陷预测研究的出版趋势如何?
- RQ2原始研究中实验报告的质量在完整性和一致性方面如何?
- RQ3哪些无监督学习算法和模型族在软件缺陷预测中表现出更优的预测性能?
- RQ4无监督模型在项目内和跨项目缺陷预测中的性能与有监督模型相比如何?
- RQ5数据集特征如何影响无监督模型的预测性能?
主要发现
- 无监督学习模型,特别是模糊C-均值(FCM)和模糊SOM(FSOM),在14种无监督模型族中表现最佳。
- FCM和FSOM在项目内与跨项目缺陷预测场景中,性能均与有监督模型相当。
- 约11%(262项,共2,456项)的报告结果存在内部不一致,表明实验计算或报告中存在错误。
- 约33%(823项,共2,456项)的结果缺乏足够细节以供验证,严重限制了可重现性与可靠性。
- 尽管性能表现强劲,但该领域仍存在实验报告标准差的问题,许多研究未能披露关键的方法学细节。
- 元分析证实,无监督模型是有监督模型的可行替代方案,尤其在标注数据稀缺的情况下,但前提是必须提升报告的严谨性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。