[论文解读] A psychophysics approach for quantitative comparison of interpretable computer vision models
本文提出一种基于心理物理学的方法,通过测量人类在图像标注任务中的表现,对可解释性计算机视觉模型进行定量评估。结果表明,与往往无法反映人类感知、可能扭曲模型实用性的纯机器指标相比,人机协同评估能提供更清晰、更具代表性的可解释性方法排名。
The field of transparent Machine Learning (ML) has contributed many novel methods aiming at better interpretability for computer vision and ML models in general. But how useful the explanations provided by transparent ML methods are for humans remains difficult to assess. Most studies evaluate interpretability in qualitative comparisons, they use experimental paradigms that do not allow for direct comparisons amongst methods or they report only offline experiments with no humans in the loop. While there are clear advantages of evaluations with no humans in the loop, such as scalability, reproducibility and less algorithmic bias than with humans in the loop, these metrics are limited in their usefulness if we do not understand how they relate to other metrics that take human cognition into account. Here we investigate the quality of interpretable computer vision algorithms using techniques from psychophysics. In crowdsourced annotation tasks we study the impact of different interpretability approaches on annotation accuracy and task time. In order to relate these findings to quality measures for interpretability without humans in the loop we compare quality metrics with and without humans in the loop. Our results demonstrate that psychophysical experiments allow for robust quality assessment of transparency in machine learning. Interestingly the quality metrics computed without humans in the loop did not provide a consistent ranking of interpretability methods nor were they representative for how useful an explanation was for humans. These findings highlight the potential of methods from classical psychophysics for modern machine learning applications. We hope that our results provide convincing arguments for evaluating interpretability in its natural habitat, human-ML interaction, if the goal is to obtain an authentic assessment of interpretability.
研究动机与目标
- 开发一种基于心理物理学方法的定量、以人类为中心的可解释性计算机视觉模型评估框架。
- 解决当前不同方法之间缺乏标准化、可比较的可解释性质量评估指标的问题。
- 探究基于机器的、无人员参与的(NHIL)可解释性指标是否能可靠反映人类感知到的可解释性。
- 评估可解释性方法对人类标注准确率、任务时间及算法偏见的影响。
- 验证心理物理学实验能否为人类-机器学习交互中的可解释性提供稳健且真实的评估。
提出的方法
- 开展众包心理物理学实验,让人类标注员在有和没有模型解释的情况下执行图像标注任务。
- 使用多种可解释性方法(如 Grad-CAM、Guided Backprop)生成的显著性图作为视觉解释,以引导人类决策。
- 通过不同解释条件下的标注准确率和任务时间来测量人类表现。
- 将人机协同(HIL)指标与无人员参与(NHIL)指标(如 L2 距离、AUC、Jaccard 相似度)进行比较。
- 通过测量当模型预测错误时,人类与模型预测之间的重叠程度,分析算法偏见。
- 采用标准化的心理物理学实验设计,以确保不同方法之间的可重复性和可比性。
实验结果
研究问题
- RQ1在受控的心理物理学实验中,不同可解释性方法如何影响人类标注准确率和任务时间?
- RQ2NHIL 指标(如 L2、AUC、Jaccard)与 HIL 性能指标(准确率、时间)的相关性有多大?
- RQ3与人类评估相比,NHIL 指标是否能对可解释性方法提供一致且具代表性的排名?
- RQ4解释质量如何影响算法偏见,即人类在未加批判地跟随错误模型预测时的表现?
- RQ5心理物理学实验能否作为人类-机器学习交互中可解释性评估的可靠、标准化基准?
主要发现
- 人机协同评估揭示了可解释性方法的清晰排名,其中 Guided Backprop 在掩码尺寸为 6% 至 19% 时表现出最高的标注准确率。
- NHIL 指标(如 L2 距离、AUC、Jaccard 相似度)在不同阈值下未能对可解释性方法产生一致的排名。
- NHIL 指标与 HIL 性能之间无显著相关性,表明基于机器的指标无法可靠反映人类感知到的可解释性。
- 尽管 Guided Backprop 方法在提升人类准确率方面最有效,但也导致了最高的算法偏见,因为标注员更常复制模型的错误预测。
- 心理物理学实验提供了稳健、稳定且可解释的可解释性质量排名,凸显其作为评估金标准的价值。
- 本研究结论认为,NHIL 指标不能代表以人类为中心的可解释性,不应作为评估模型透明度的唯一依据。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。