[Paper Review] A psychophysics approach for quantitative comparison of interpretable computer vision models
This paper proposes a psychophysics-based approach to quantitatively evaluate interpretable computer vision models by measuring human performance in image annotation tasks. It demonstrates that human-in-the-loop evaluations yield clearer, more representative rankings of interpretability methods than machine-only metrics, which often fail to correlate with human perception and can misrepresent model utility.
The field of transparent Machine Learning (ML) has contributed many novel methods aiming at better interpretability for computer vision and ML models in general. But how useful the explanations provided by transparent ML methods are for humans remains difficult to assess. Most studies evaluate interpretability in qualitative comparisons, they use experimental paradigms that do not allow for direct comparisons amongst methods or they report only offline experiments with no humans in the loop. While there are clear advantages of evaluations with no humans in the loop, such as scalability, reproducibility and less algorithmic bias than with humans in the loop, these metrics are limited in their usefulness if we do not understand how they relate to other metrics that take human cognition into account. Here we investigate the quality of interpretable computer vision algorithms using techniques from psychophysics. In crowdsourced annotation tasks we study the impact of different interpretability approaches on annotation accuracy and task time. In order to relate these findings to quality measures for interpretability without humans in the loop we compare quality metrics with and without humans in the loop. Our results demonstrate that psychophysical experiments allow for robust quality assessment of transparency in machine learning. Interestingly the quality metrics computed without humans in the loop did not provide a consistent ranking of interpretability methods nor were they representative for how useful an explanation was for humans. These findings highlight the potential of methods from classical psychophysics for modern machine learning applications. We hope that our results provide convincing arguments for evaluating interpretability in its natural habitat, human-ML interaction, if the goal is to obtain an authentic assessment of interpretability.
Motivation & Objective
- To develop a quantitative, human-centered evaluation framework for interpretable computer vision models using psychophysical methods.
- To address the lack of standardized, comparable metrics for evaluating interpretability quality across different methods.
- To investigate whether machine-based, no-humans-in-the-loop (NHIL) interpretability metrics reliably reflect human-perceived interpretability.
- To assess the impact of interpretability methods on human annotation accuracy, task time, and algorithmic bias.
- To validate that psychophysical experiments provide a robust and authentic assessment of interpretability in human-ML interaction.
Proposed method
- Conducted crowdsourced psychophysical experiments where human annotators performed image annotation tasks with and without model explanations.
- Used saliency maps from multiple interpretability methods (e.g., Grad-CAM, Guided Backprop) as visual explanations to guide human decisions.
- Measured human performance via annotation accuracy and task time under different explanation conditions.
- Compared human-in-the-loop (HIL) metrics with no-humans-in-the-loop (NHIL) metrics such as L2 distance, AUC, and Jaccard similarity.
- Analyzed algorithmic bias by measuring overlap between human and model predictions when the model was incorrect.
- Applied standardized psychophysical experimental design to ensure reproducibility and comparability across methods.
Experimental results
Research questions
- RQ1How do different interpretability methods affect human annotation accuracy and task time in a controlled psychophysical experiment?
- RQ2To what extent do NHIL metrics (e.g., L2, AUC, Jaccard) correlate with HIL performance metrics (accuracy, time)?
- RQ3Do NHIL metrics provide a consistent and representative ranking of interpretability methods compared to human evaluations?
- RQ4How does the quality of explanations influence algorithmic bias, where humans uncritically follow incorrect model predictions?
- RQ5Can psychophysical experiments serve as a reliable, standardized benchmark for evaluating interpretability in human-ML interaction?
Key findings
- Human-in-the-loop evaluations revealed a clear ranking of interpretability methods, with Guided Backprop showing the highest annotation accuracy for mask sizes of 6% to 19%.
- NHIL metrics such as L2 distance, AUC, and Jaccard similarity failed to produce a consistent ranking of interpretability methods across different thresholds.
- There was no significant correlation between NHIL metrics and HIL performance, indicating that machine-based metrics do not reliably reflect human-perceived interpretability.
- The Guided Backprop method, while most helpful for human accuracy, also led to the highest level of algorithmic bias, as annotators more often replicated the model’s incorrect predictions.
- Psychophysical experiments provided robust, stable, and interpretable rankings of interpretability quality, highlighting their value as a gold standard for evaluation.
- The study concludes that NHIL metrics are not representative of human-centered interpretability and should not be used as the sole basis for evaluating model transparency.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.