[Paper Review] Quantifying Interpretability and Trust in Machine Learning Systems
The paper proposes quantitative metrics to measure interpretability and trust in ML decisions, and demonstrates their use through crowdsourcing experiments and tasks showing interpretability can boost human productivity while revealing biased trust.
Decisions by Machine Learning (ML) models have become ubiquitous. Trusting these decisions requires understanding how algorithms take them. Hence interpretability methods for ML are an active focus of research. A central problem in this context is that both the quality of interpretability methods as well as trust in ML predictions are difficult to measure. Yet evaluations, comparisons and improvements of trust and interpretability require quantifiable measures. Here we propose a quantitative measure for the quality of interpretability methods. Based on that we derive a quantitative measure of trust in ML decisions. Building on previous work we propose to measure intuitive understanding of algorithmic decisions using the information transfer rate at which humans replicate ML model predictions. We provide empirical evidence from crowdsourcing experiments that the proposed metric robustly differentiates interpretability methods. The proposed metric also demonstrates the value of interpretability for ML assisted human decision making: in our experiments providing explanations more than doubled productivity in annotation tasks. However unbiased human judgement is critical for doctors, judges, policy makers and others. Here we derive a trust metric that identifies when human decisions are overly biased towards ML predictions. Our results complement existing qualitative work on trust and interpretability by quantifiable measures that can serve as objectives for further improving methods in this field of research.
Motivation & Objective
- Motivate the need for measurable interpretability and trust in ML decisions.
- Propose a quantitative metric for the quality of interpretability methods.
- Derive a trust metric that identifies biased human decisions influenced by ML predictions.
Proposed method
- Define a metric based on information transfer rate capturing how well humans replicate ML model predictions under explanations.
- Use crowdsourcing experiments to evaluate how different interpretability methods affect human understanding and performance.
- Demonstrate that providing explanations can more than double productivity in annotation tasks.
- Derive a trust metric to detect when human judgments are overly biased toward ML predictions.
Experimental results
Research questions
- RQ1Can a quantitative measure reliably differentiate interpretable explanations from less interpretable ones?
- RQ2Does interpretability improve human decision-making efficiency and accuracy in ML-assisted tasks?
- RQ3Under what conditions do humans exhibit biased trust toward ML predictions, and how can this be quantified?
Key findings
- The proposed information transfer rate metric robustly differentiates interpretability methods in crowdsourcing studies.
- Explanations significantly increase productivity in annotation tasks (more than doubling).
- A derived trust metric identifies when human decisions are overly biased toward ML predictions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.