[论文解读] A Human-Grounded Evaluation of SHAP for Alert Processing
本研究通过159名参与者参与的人工实证实验,评估了SHAP解释在警报处理任务中的表现。尽管任务表现未出现显著提升,但SHAP值通过引导注意力聚焦于显著特征,影响了决策过程,表明其价值在于引导注意力,而非单纯提升准确性。
In the past years, many new explanation methods have been proposed to achieve interpretability of machine learning predictions. However, the utility of these methods in practical applications has not been researched extensively. In this paper we present the results of a human-grounded evaluation of SHAP, an explanation method that has been well-received in the XAI and related communities. In particular, we study whether this local model-agnostic explanation method can be useful for real human domain experts to assess the correctness of positive predictions, i.e. alerts generated by a classifier. We performed experimentation with three different groups of participants (159 in total), who had basic knowledge of explainable machine learning. We performed a qualitative analysis of recorded reflections of experiment participants performing alert processing with and without SHAP information. The results suggest that the SHAP explanations do impact the decision-making process, although the model's confidence score remains to be a leading source of evidence. We statistically test whether there is a significant difference in task utility metrics between tasks for which an explanation was available and tasks in which it was not provided. As opposed to common intuitions, we did not find a significant difference in alert processing performance when a SHAP explanation is available compared to when it is not.
研究动机与目标
- 评估SHAP解释在涉及人类专家的实际决策任务中的现实应用价值。
- 评估SHAP是否提升了警报处理任务中的任务有效性、效率及心理负担。
- 通过分析决策过程中的书面反思,探究SHAP如何影响人类推理过程。
- 确定SHAP解释是否与仅依赖模型置信度相比,导致可测量的性能差异。
提出的方法
- 在159名参与者中开展两项受控实验,使用学生表现数据集完成简化的警报处理任务。
- 参与者以交替条件在有和无SHAP解释的情况下,判断预测为真阳性或假阳性。
- 通过标准化指标和任务后问卷,测量任务有效性、效率及心理负担。
- 使用自然语言处理技术分析书面推理,比较SHAP与NoSHAP条件下特征提及频率的差异。
- 采用统计检验(如t检验、ANOVA)比较不同条件下的表现,并通过事后等价性检验评估无显著差异。
- 通过0.2个百分点的术语比例差异阈值,识别SHAP值导致特征提及频率增加的案例。
实验结果
研究问题
- RQ1与仅依赖置信度分数相比,提供SHAP解释是否显著提升警报处理任务的有效性?
- RQ2SHAP解释的可用性如何影响人类决策中的任务效率和心理负担?
- RQ3在警报评估过程中,SHAP解释在多大程度上影响了领域专家的推理过程?
- RQ4是否存在特定特征或实例,使得SHAP解释能带来更一致或更准确的推理?
主要发现
- 在有和无SHAP解释的任务中,任务有效性、效率及心理负担均未发现统计学上的显著差异。
- 参与者主要依赖模型的置信度分数作为主要证据来源,这可能具有误导性。
- 较大的SHAP值(绝对值在0.07至0.11之间)与书面推理中对特定特征值的关注度提升相关。
- 在14个实例中的5个,SHAP条件下对特征值的讨论频率更高,表明其对推理过程有可测量的影响。
- 零假设未能被拒绝,很可能是由于效应量较小,而非数据不足,表明SHAP在孤立使用时效用有限。
- SHAP解释并未导致参与者放弃先前考虑的特征值,而是突出了此前被忽略的特征。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。