[论文解读] Machine Explanations and Human Understanding
本文提出了一套形式化的理论框架,将机器解释与人类在人机决策中的理解联系起来,证明只有当解释基于与任务相关的特定人类直觉时,才能提升对模型行为的理解,而非对任务决策或错误的理解。关键贡献在于证明了有效的人机协作需要显式建模人类直觉,并通过受控的人类受试研究验证了这一点,显示若缺乏此类直觉,人类会过度依赖AI。
Explanations are hypothesized to improve human understanding of machine learning models and achieve a variety of desirable outcomes, ranging from model debugging to enhancing human decision making. However, empirical studies have found mixed and even negative results. An open question, therefore, is under what conditions explanations can improve human understanding and in what way. Using adapted causal diagrams, we provide a formal characterization of the interplay between machine explanations and human understanding, and show how human intuitions play a central role in enabling human understanding. Specifically, we identify three core concepts of interest that cover all existing quantitative measures of understanding in the context of human-AI decision making: task decision boundary, model decision boundary, and model error. Our key result is that without assumptions about task-specific intuitions, explanations may potentially improve human understanding of model decision boundary, but they cannot improve human understanding of task decision boundary or model error. To achieve complementary human-AI performance, we articulate possible ways on how explanations need to work with human intuitions. For instance, human intuitions about the relevance of features (e.g., education is more important than age in predicting a person's income) can be critical in detecting model error. We validate the importance of human intuitions in shaping the outcome of machine explanations with empirical human-subject studies. Overall, our work provides a general framework along with actionable implications for future algorithmic development and empirical experiments of machine explanations.
研究动机与目标
- 通过识别解释提升人类理解的条件,解决关于机器解释的矛盾实证结果。
- 形式化理解的三个核心概念之间的区别:任务决策边界、模型决策边界和模型错误。
- 证明在缺乏对人类直觉的假设时,机器解释无法提升人类对任务或错误边界的理解。
- 通过Wizard-of-Oz人类受试研究,受控操纵人类直觉,验证理论主张。
- 倡导在研究设计和解释系统开发中明确表达人类直觉。
提出的方法
- 开发适配的因果图,以建模人类对任务/模型边界的近似与模型解释之间的关系。
- 引入一个正式的 $\text{show}$ 算子,表示干预(如揭示模型预测)如何塑造人类理解。
- 根据任务背景,将理解分为两类:模仿(理解模型)与发现(理解任务)。
- 提出解释必须与人类直觉(如特征重要性或相关性方向)对齐,才能提升对任务层面决策的理解。
- 采用Wizard-of-Oz实验设置,隔离并控制人类直觉,比较有和无假设直觉的组别之间的共识率。
- 使用独立样本和配对t检验,比较不同条件下人类-AI决策的一致性,衡量解释的一致性与对齐度。

实验结果
研究问题
- RQ1在什么条件下,机器解释可以提升人类对机器学习模型的理解?
- RQ2尽管有理论预期,为何关于解释的实证研究仍报告混合或负面结果?
- RQ3特定任务的人类直觉如何影响人机决策中机器解释的有效性?
- RQ4在缺乏人类直觉假设的情况下,解释能否提升对任务决策边界或模型错误的理解?
- RQ5缺乏人类直觉在多大程度上导致对AI预测的过度依赖?
主要发现
- 在缺乏对人类直觉的假设时,机器解释只能提升对模型决策边界的理解,而无法提升对任务决策边界或模型错误的理解。
- 在匿名组(无直觉)中,参与者与AI预测一致的比例为70.64%,显著高于常规组(有直觉)的54.32%,表明在缺乏直觉时存在过度依赖。
- 当直觉存在时,对一致解释对(AB)的共识率为90.71%,而对不一致解释对(CD)仅为25.00%,表明与正确直觉高度对齐。
- 当解释与人类直觉一致时,一致性显著更高:AB组为90.71%,EF组为22.14%,p值 < 0.001。
- t检验结果证实,不同条件下共识率的差异具有统计显著性(p < 0.001),验证了直觉在解释有效性中的作用。
- 研究证实,若无特定任务的人类直觉,人类与AI协同表现(人类+AI > 人类且 > AI)将不可能实现,与理论框架预测一致。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。