[Paper Review] Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead
The paper argues that high-stakes decisions should rely on inherently interpretable models rather than explanations of black-box models, and discusses why explainable ML is flawed and how interpretable ML can be developed and governed.
Black box machine learning models are currently being used for high stakes decision-making throughout society, causing problems throughout healthcare, criminal justice, and in other domains. People have hoped that creating methods for explaining these black box models will alleviate some of these problems, but trying to extit{explain} black box models, rather than creating models that are extit{interpretable} in the first place, is likely to perpetuate bad practices and can potentially cause catastrophic harm to society. There is a way forward -- it is to design models that are inherently interpretable. This manuscript clarifies the chasm between explaining black boxes and using inherently interpretable models, outlines several key reasons why explainable black boxes should be avoided in high-stakes decisions, identifies challenges to interpretable machine learning, and provides several example applications where interpretable models could potentially replace black box models in criminal justice, healthcare, and computer vision.
Motivation & Objective
- Clarify the limitations and risks of explaining black-box ML in high-stakes domains.
- Advocate for using intrinsically interpretable models that provide faithful explanations by design.
- Identify domain-specific interpretability constraints and benefits across fields like criminal justice and healthcare.
- Discuss governance, policy implications, and practical challenges in adopting interpretable ML.
Proposed method
- Critically compare explainable ML with inherently interpretable models and discuss fidelity issues between explanations and original models.
- Present the CORELS interpretable rule-list approach as an example of a sparse, interpretable model with competitive accuracy.
- Outline architectural and optimization challenges in constructing interpretable models under domain-specific constraints.
- Provide qualitative comparisons between black-box models and interpretable alternatives in real-world domains.
Experimental results
Research questions
- RQ1Can inherently interpretable models achieve comparable accuracy to black-box models in high-stakes tasks?
- RQ2What are the major practical and governance barriers to adopting interpretable ML in sensitive domains?
- RQ3How can interpretable models be constructed to satisfy domain knowledge such as sparsity, additivity, monotonicity, or causality?
- RQ4What are the limitations of post-hoc explanations for black-box models in high-stakes decisions?
- RQ5What policy proposals could incentivize the shift from black-box to interpretable models?
Key findings
- Explainable post-hoc methods often produce faithful approximate explanations only with caveats and can be misleading.
- Interpretable models can achieve comparable predictive performance to black-box models in structured data settings when meaningful features are used.
- There are significant governance and business-model barriers to deploying proprietary black-box models in high-stakes contexts.
- Interpretable models like CORELS demonstrate that simple, transparent rules can rival complex models in certain recidivism prediction tasks.
- Policy and regulation considerations could mandate the use of interpretable models when they offer similar performance to black-box alternatives.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.