Skip to main content
QUICK REVIEW

[Paper Review] Exploring ML testing in practice -- Lessons learned from an interactive rapid review with Axis Communications

Qunying Song, Markus Borg|arXiv (Cornell University)|Mar 30, 2022
Software Engineering Research4 citations
TL;DR

This study conducts an interactive rapid review (IRR) with industry practitioners from Axis Communications and researchers to explore machine learning (ML) testing challenges, particularly data testing in computer vision systems. The collaboration produced a refined taxonomy, identified 12 key challenges, and delivered nine technological rules for data testing—offering practical, adaptable insights despite no perfect match to industry needs.

ABSTRACT

There is a growing interest in industry and academia in machine learning (ML) testing. We believe that industry and academia need to learn together to produce rigorous and relevant knowledge. In this study, we initiate a collaboration between stakeholders from one case company, one research institute, and one university. To establish a common view of the problem domain, we applied an interactive rapid review of the state of the art. Four researchers from Lund University and RISE Research Institutes and four practitioners from Axis Communications reviewed a set of 180 primary studies on ML testing. We developed a taxonomy for the communication around ML testing challenges and results and identified a list of 12 review questions relevant for Axis Communications. The three most important questions (data testing, metrics for assessment, and test generation) were mapped to the literature, and an in-depth analysis of the 35 primary studies matching the most important question (data testing) was made. A final set of the five best matches were analysed and we reflect on the criteria for applicability and relevance for the industry. The taxonomies are helpful for communication but not final. Furthermore, there was no perfect match to the case company's investigated review question (data testing). However, we extracted relevant approaches from the five studies on a conceptual level to support later context-specific improvements. We found the interactive rapid review approach useful for triggering and aligning communication between the different stakeholders.

Motivation & Objective

  • To initiate a collaborative research effort between academia and industry on ML testing challenges in real-world applications.
  • To align terminology and research priorities between practitioners at Axis Communications and academic researchers.
  • To identify and map relevant research on ML testing to industry-specific challenges, especially data testing in computer vision.
  • To extract transferable technological rules and applicability criteria from academic literature for industrial use.
  • To support future joint research and MSc thesis projects focused on data testing in safety-critical ML systems.

Proposed method

  • Conducted an interactive rapid review (IRR) involving 8 stakeholders: 4 researchers and 4 practitioners from Axis Communications.
  • Reviewed 180 primary studies on ML testing using a structured, collaborative process with iterative alignment and discussion.
  • Developed and extended a taxonomy (SERP taxonomy) based on existing software testing and AI quality frameworks to classify challenges and solutions.
  • Prioritized 12 practical challenges from Axis, with the top three (data testing, metrics, test generation) mapped to relevant literature.
  • Performed an in-depth analysis of 35 studies on 'how to test the dataset', extracting nine technological rules and context factors.
  • Evaluated relevance and applicability of selected studies using criteria such as conceptual fit, technical feasibility, and domain alignment.

Experimental results

Research questions

  • RQ1What are the most relevant and applicable research solutions for data testing in ML-based computer vision systems from industry’s perspective?
  • RQ2How can researchers and practitioners collaboratively identify and prioritize key challenges in ML testing?
  • RQ3What technological rules and criteria can be extracted from academic literature to guide industrial application of ML testing?
  • RQ4How can a taxonomy support communication and alignment between academic research and industrial practice in AI quality?
  • RQ5What are the key context factors that affect the applicability of ML testing techniques in real-world systems?

Key findings

  • No perfect match was found between existing academic research and Axis Communications’ specific data testing challenges, indicating a research gap in practical, context-specific data testing for industrial computer vision systems.
  • Nine technological rules for data testing were extracted from five high-potential studies, offering conceptual frameworks that can be adapted to industrial needs.
  • The interactive rapid review process successfully facilitated shared understanding, aligned terminology, and built trust between academic and industrial stakeholders.
  • The SERP taxonomy, extended with software testing and AI quality dimensions, proved useful for organizing and navigating the complex landscape of ML testing research.
  • Practitioners found the surprise adequacy technique particularly relevant and have already begun applying it in data collection processes.
  • The study highlights the importance of data quality assurance as central to AI quality, especially in safety-critical domains like automotive, avionics, and healthcare.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.