Skip to main content
QUICK REVIEW

[Paper Review] A Systematic Mapping Study on Testing of Machine Learning Programs

Salman Sherin, Muhammad Uzair Khan|arXiv (Cornell University)|Jul 11, 2019
Software Testing and Debugging Techniques47 references8 citations
TL;DR

This systematic mapping study analyzes 37 relevant papers from 1,654 identified sources to map the state of testing in machine learning programs. It identifies key trends, classifies testing techniques by ML type and approach, and highlights gaps in empirical evidence, tool availability, and non-functional testing—particularly for reinforcement learning systems.

ABSTRACT

We aim to conduct a systematic mapping in the area of testing ML programs. We identify, analyze and classify the existing literature to provide an overview of the area. We followed well-established guidelines of systematic mapping to develop a systematic protocol to identify and review the existing literature. We formulate three sets of research questions, define inclusion and exclusion criteria and systematically identify themes for the classification of existing techniques. We also report the quality of the published works using established assessment criteria. we finally selected 37 papers out of 1654 based on our selection criteria up to January 2019. We analyze trends such as contribution facet, research facet, test approach, type of ML and the kind of testing with several other attributes. We also discuss the empirical evidence and reporting quality of selected papers. The data from the study is made publicly available for other researchers and practitioners. We present an overview of the area by answering several research questions. The area is growing rapidly, however, there is lack of enough empirical evidence to compare and assess the effectiveness of the techniques. More publicly available tools are required for use of practitioners and researchers. Further attention is needed on non-functional testing and testing of ML programs using reinforcement learning. We believe that this study can help researchers and practitioners to obtain an overview of the area and identify several sub-areas where more research is required

Motivation & Objective

  • To provide a comprehensive overview of the current state of research on testing machine learning programs.
  • To identify and classify existing testing techniques based on ML type, test approach, and research focus.
  • To assess the quality of published works in terms of empirical evidence and reporting standards.
  • To highlight under-researched areas such as non-functional testing and reinforcement learning-based ML systems.
  • To support researchers and practitioners by making data publicly available and identifying future research directions.

Proposed method

  • Conducted a systematic mapping study using established guidelines for literature review in software engineering.
  • Defined inclusion and exclusion criteria to filter 1,654 initial papers down to 37 relevant studies up to January 2019.
  • Formulated three sets of research questions to guide classification and analysis of the selected literature.
  • Classified papers based on attributes including contribution facet, research facet, test approach, ML type, and testing type.
  • Evaluated the quality of selected papers using established assessment criteria for empirical reporting and reproducibility.
  • Publicly released the collected data to support future research and tool development in ML testing.

Experimental results

Research questions

  • RQ1What are the dominant research themes and trends in testing machine learning programs?
  • RQ2How are testing techniques categorized by machine learning type and test approach?
  • RQ3What is the quality of empirical evidence and reporting in published ML testing studies?
  • RQ4Which areas of ML testing remain under-investigated, particularly in non-functional testing and reinforcement learning?
  • RQ5What are the key gaps in tool availability and reproducibility for ML testing research and practice?

Key findings

  • The field of ML program testing is growing rapidly, but lacks sufficient empirical evidence to compare or assess the effectiveness of different testing techniques.
  • Only 37 out of 1,654 initially identified papers met the inclusion criteria, indicating a high level of noise and irrelevance in the literature.
  • There is a notable lack of publicly available tools to support practitioners and researchers in testing ML systems.
  • Non-functional testing and testing of reinforcement learning-based ML programs remain under-investigated and represent significant research gaps.
  • The reporting quality of selected papers is inconsistent, with limited use of standardized evaluation metrics and reproducible experimental setups.
  • The study's data set is publicly available to support future research, tool development, and benchmarking in ML testing.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.