Skip to main content
QUICK REVIEW

[Paper Review] Machine Learning for Software Engineering: A Tertiary Study

Zoe Kotti, Rafaila Galanopoulou|arXiv (Cornell University)|Nov 17, 2022
Software Engineering Research176 references4 citations
TL;DR

This tertiary study synthesizes 83 reviews on machine learning (ML) in software engineering (SE), analyzing 6,117 primary studies to identify trends, challenges, and research gaps. It reveals that ML is most applied in software quality and testing, with underrepresentation in human-centered SE areas, and proposes key actions like improving data pipelines, adopting incremental learning, and enhancing industrial collaboration to advance ML for SE applications.

ABSTRACT

Machine learning (ML) techniques increase the effectiveness of software engineering (SE) lifecycle activities. We systematically collected, quality-assessed, summarized, and categorized 83 reviews in ML for SE published between 2009-2022, covering 6,117 primary studies. The SE areas most tackled with ML are software quality and testing, while human-centered areas appear more challenging for ML. We propose a number of ML for SE research challenges and actions including: conducting further empirical validation and industrial studies on ML; reconsidering deficient SE methods; documenting and automating data collection and pipeline processes; reexamining how industrial practitioners distribute their proprietary data; and implementing incremental ML approaches.

Motivation & Objective

  • To provide a systematic, evidence-based overview of the current state of machine learning (ML) applications in software engineering (SE) by synthesizing existing secondary literature.
  • To identify the most frequently targeted SE areas, ML techniques, and research gaps in ML for SE through a large-scale review of reviews.
  • To highlight challenges such as data quality, lack of industrial relevance, and insufficient empirical validation in ML-based SE research.
  • To propose actionable research and industrial collaboration strategies to improve the robustness, scalability, and real-world applicability of ML in SE.
  • To differentiate ML for SE from the related field of SE for ML, ensuring clarity in research scope and direction.

Proposed method

  • Systematically collected and quality-assessed 83 secondary reviews on ML for SE published between 2009 and 2022, following PRISMA and tertiary study guidelines.
  • Classified reviews using the SWEBOK knowledge areas (KAs) and subareas to map ML applications across the SE lifecycle.
  • Applied a four-axis ML classification scheme—based on learning type, task type, model type, and data type—to categorize ML techniques used in the reviews.
  • Extracted and analyzed SE tasks, research topics, and ML techniques from each review through manual, iterative coding and consensus among authors.
  • Evaluated data collection practices, model generalizability, and industrial relevance by assessing documentation, pipeline automation, and data sharing practices in the reviewed studies.
  • Synthesized recommendations for researchers and practitioners based on findings, including the need for empirical validation, industrial case studies, and improved data management.

Experimental results

Research questions

  • RQ1Which software engineering knowledge areas and tasks are most frequently targeted by machine learning applications, according to secondary reviews?
  • RQ2What are the dominant types of machine learning techniques (e.g., supervised, online learning) used in SE, and how do they compare in terms of adoption and generalizability?
  • RQ3What are the major barriers to the industrial adoption of ML in SE, particularly concerning data quality, pipeline automation, and proprietary data sharing?
  • RQ4How do deficiencies in SE methods and data collection practices affect the reliability and scalability of ML models in software engineering?
  • RQ5What research and collaboration strategies can bridge the gap between academic ML research and industrial SE practice?

Key findings

  • The majority of ML for SE research focuses on software quality and testing, with significantly less attention given to human-centered areas like software requirements and professional practice due to data subjectivity and labeling challenges.
  • Supervised learning and classification tasks dominate, with model-based learning preferred over online/incremental learning despite its demonstrated benefits for real-time and evolving systems.
  • A critical gap exists in data pipeline transparency, with many studies lacking documented data collection and non-automated processes, undermining reproducibility and scalability.
  • Industrial relevance is hampered by the absence of proprietary data sharing, limiting model performance and generalization in real-world SE environments.
  • Outdated and inadequate SE methods obstruct the development of robust ML models, highlighting the need for methodological reevaluation and improved hyper-parameter tuning and class imbalance handling.
  • Hybrid, ensemble, and incremental ML techniques represent underexplored but promising research avenues for improving model reliability and adaptability in SE contexts.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.