[論文レビュー] A Systematic Review of Unsupervised Learning Techniques for Software Defect Prediction
本システマティックレビューは、49件の研究および2,456件の実験結果をメタアナリシスした結果、教師なし学習手法がソフトウェア欠陥予測において教師あり手法と同等の性能を示すことを評価している。Fuzzy C-Means (FCM) や Fuzzy SOMs (FSOMs) といった教師なしモデルは、報告品質と実験の一貫性に著しい問題が広がっているため、信頼性が損なわれている。
Background: Unsupervised machine learners have been increasingly applied to software defect prediction. It is an approach that may be valuable for software practitioners because it reduces the need for labeled training data. Objective: Investigate the use and performance of unsupervised learning techniques in software defect prediction. Method: We conducted a systematic literature review that identified 49 studies containing 2456 individual experimental results, which satisfied our inclusion criteria published between January 2000 and March 2018. In order to compare prediction performance across these studies in a consistent way, we (re-)computed the confusion matrices and employed the Matthews Correlation Coefficient (MCC) as our main performance measure. Results: Our meta-analysis shows that unsupervised models are comparable with supervised models for both within-project and cross-project prediction. Among the 14 families of unsupervised model, Fuzzy CMeans (FCM) and Fuzzy SOMs (FSOMs) perform best. In addition, where we were able to check, we found that almost 11% (262/2456) of published results (contained in 16 papers) were internally inconsistent and a further 33% (823/2456) provided insufficient details for us to check. Conclusion: Although many factors impact the performance of a classifier, e.g., dataset characteristics, broadly speaking, unsupervised classifiers do not seem to perform worse than the supervised classifiers in our review. However, we note a worrying prevalence of (i) demonstrably erroneous experimental results, (ii) undemanding benchmarks and (iii) incomplete reporting. We therefore encourage researchers to be comprehensive in their reporting.
研究の動機と目的
- 教師あり手法と比較して、教師なし学習手法のソフトウェア欠陥予測における性能を評価すること。
- プロジェクト内およびプロジェクト間の両設定において、欠陥予測に最も効果的な教師なしモデルファミリーを特定すること。
- 報告の完全性と一貫性に注目し、教師なし欠陥予測研究における実験報告の質を評価すること。
- 実務家および研究者に対して、教師なし欠陥予測モデルの実用性と信頼性に関する実行可能なガイダンスを提供すること。
- 再現性と妥当性を損なう、実験設計および報告における体系的な問題を浮き彫りにすること。
提案手法
- 2000年1月から2018年3月までの期間、Web of Science、ACM、IEEE Xplore、ScienceDirect、SpringerLinkの5つの主要データベースを用いてシステマティックレビューを実施した。
- 事前に定義された含む基準を適用して、教師なし欠陥予測に関する49件の一次研究を特定し、合計2,456件の個別の実験結果を取得した。
- 報告された性能指標から再計算された混同行列を用いて、研究間での一貫性を確保し、公平な比較を可能にした。
- メタアナリシスの主な性能指標として、研究間比較を可能にするためのスレーシャー相関係数(MCC)を採用した。
- 文献計測および品質評価分析を実施し、報告の完全性、一貫性、研究トレンドを評価した。
- モデルを14のファミリーおよび6つのクラスタラベル付け手法に分類し、アルゴリズムファミリーごとの性能を比較した。
実験結果
リサーチクエスチョン
- RQ12000年から2018年の間、教師なしソフトウェア欠陥予測研究における出版トレンドはどのようなものか?
- RQ2一次研究全体を通じて、報告の完全性と一貫性という観点から、実験報告の質はどの程度か?
- RQ3どの教師なし学習アルゴリズムおよびモデルファミリーがソフトウェア欠陥予測において優れた予測性能を示すか?
- RQ4教師ありモデルと比較して、教師なしモデルはプロジェクト内およびプロジェクト間の欠陥予測において、性能でどのように差がつくか?
- RQ5データセットの特性は、教師なしモデルの予測性能にどのように影響を与えるか?
主な発見
- 教師なし学習モデルのうち、特にFuzzy C-Means (FCM) および Fuzzy SOMs (FSOMs) が14の教師なしモデルファミリーの中で最高の性能を示した。
- FCMおよびFSOMsは、プロジェクト内およびプロジェクト間の両方の欠陥予測シナリオにおいて、教師ありモデルと同等の性能を達成した。
- 約11%(2,456件中262件)の報告結果に内部的一致性が欠けており、実験計算または報告に誤りがある可能性を示している。
- 約33%(2,456件中823件)の結果は、検証が可能な十分な詳細が欠落しており、再現性と信頼性が著しく制限されている。
- 強力な性能を示しているにもかかわらず、分野全体が報告基準が低いことが問題となっており、多くの研究が重要な手法論的詳細を明らかにしていない。
- メタアナリシスにより、教師ありモデルの代替として教師なしモデルが有効であることが確認されたが、特にラベル付きデータが限られる状況では、報告の厳密さを向上させる必要がある。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。