Skip to main content
QUICK REVIEW

[Paper Review] Assuring the Machine Learning Lifecycle: Desiderata, Methods, and Challenges

Rob Ashmore, Radu Călinescu|arXiv (Cornell University)|May 10, 2019
Adversarial Robustness in Machine LearningComputer Science160 references87 citations
TL;DR

A comprehensive survey of assurance for the ML lifecycle, detailing desiderata and methods for data management, model learning, verification, and deployment, and outlining open challenges.

ABSTRACT

Machine learning has evolved into an enabling technology for a wide range of highly successful applications. The potential for this success to continue and accelerate has placed machine learning (ML) at the top of research, economic and political agendas. Such unprecedented interest is fuelled by a vision of ML applicability extending to healthcare, transportation, defence and other domains of great societal importance. Achieving this vision requires the use of ML in safety-critical applications that demand levels of assurance beyond those needed for current ML applications. Our paper provides a comprehensive survey of the state-of-the-art in the assurance of ML, i.e. in the generation of evidence that ML is sufficiently safe for its intended use. The survey covers the methods capable of providing such evidence at different stages of the machine learning lifecycle, i.e. of the complex, iterative process that starts with the collection of the data used to train an ML component for a system, and ends with the deployment of that component within the system. The paper begins with a systematic presentation of the ML lifecycle and its stages. We then define assurance desiderata for each stage, review existing methods that contribute to achieving these desiderata, and identify open challenges that require further research.

Motivation & Objective

  • Define the machine learning lifecycle and its four stages (Data Management, Model Learning, Model Verification, Model Deployment).
  • Identify assurance desiderata for artefacts produced at each stage to enable trustworthy ML in safety-critical systems.
  • Review existing assurance methods for each stage and discuss their assumptions, advantages, and limitations.
  • Highlight open challenges in acquiring, validating, and integrating assurance evidence across the lifecycle.

Proposed method

  • Systematically define the ML lifecycle and its stages.
  • Map assurance desiderata to artefacts produced at each stage (data sets, models, verification results).
  • Survey methods for data management, model learning, verification (including formal verification), and deployment that contribute to assurance.
  • Discuss limitations and applicability across supervised, unsupervised, and reinforcement learning.
  • Identify open challenges requiring further research and development.

Experimental results

Research questions

  • RQ1What assurance evidence is needed at each stage of the ML lifecycle to support safety-critical deployments?
  • RQ2What methods exist to generate and validate this assurance evidence across data management, model learning, verification, and deployment?
  • RQ3What are the principal challenges in achieving thorough assurance throughout the ML lifecycle?

Key findings

  • The paper provides a structured lifecycle with four stages and explains how assurance evidence should be produced at each stage.
  • It synthesizes methods across data management, model learning, verification, and deployment, along with their assumptions and limitations.
  • It discusses formal verification, test-based verification, and deployment-time monitoring as parts of assurance.
  • It identifies open challenges in data quality, completeness of input domains, leakage, bias, and adversarial robustness.
  • The survey argues for comprehensive assurance cases that integrate evidence from all lifecycle stages for safety-critical systems.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.