[Paper Review] Failure Modes in Machine Learning Systems
The paper proposes a taxonomy to classify ML system failures into intentional (adversarial) and unintentional (inherently unsafe outcomes) and discusses its applicability for practitioners and policymakers.
In the last two years, more than 200 papers have been written on how machine learning (ML) systems can fail because of adversarial attacks on the algorithms and data; this number balloons if we were to incorporate papers covering non-adversarial failure modes. The spate of papers has made it difficult for ML practitioners, let alone engineers, lawyers, and policymakers, to keep up with the attacks against and defenses of ML systems. However, as these systems become more pervasive, the need to understand how they fail, whether by the hand of an adversary or due to the inherent design of a system, will only become more pressing. In order to equip software developers, security incident responders, lawyers, and policy makers with a common vernacular to talk about this problem, we developed a framework to classify failures into "Intentional failures" where the failure is caused by an active adversary attempting to subvert the system to attain her goals; and "Unintentional failures" where the failure is because an ML system produces an inherently unsafe outcome. After developing the initial version of the taxonomy last year, we worked with security and ML teams across Microsoft, 23 external partners, standards organization, and governments to understand how stakeholders would use our framework. Throughout the paper, we attempt to highlight how machine learning failure modes are meaningfully different from traditional software failures from a technology and policy perspective.
Motivation & Objective
- Provide a common vernacular to discuss ML failures for developers, security responders, lawyers, and policymakers.
- Introduce a taxonomy that separates failures into intentional and unintentional categories.
- Show how ML failure modes differ from traditional software failures from technology and policy perspectives.
Proposed method
- Develop a framework to classify ML failures into Intentional and Unintentional categories.
- Collaborate with Microsoft, external partners, standards organizations, and governments to validate stakeholder usefulness.
- Compare ML failure modes to traditional software failures in technology and policy contexts.
Experimental results
Research questions
- RQ1How can ML system failures be categorized into actionable classes for diverse stakeholders?
- RQ2What are the defining characteristics of intentional versus unintentional ML failures?
- RQ3In what ways do ML failure modes differ from traditional software failure modes from technical and policy perspectives.
Key findings
- A taxonomy that distinguishes Intentional failures (caused by active adversaries) from Unintentional failures (inherently unsafe outcomes).
- The framework is designed to be used by software developers, incident responders, lawyers, and policymakers.
- ML failure modes exhibit meaningful differences from traditional software failures in both technology and policy dimensions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.