[论文解读] On Anomaly Interpretation via Shapley Values
本文提出了一种基于Shapley值的方法,通过定义一种新颖的特征函数,利用异常分数最小化来近似特征缺失,从而解释半监督异常检测中的异常分数。在多个数据集和模型上的实验表明,基于Shapley值的归因方法能够提升异常解释能力,并支持更有效的异常定位。
Anomaly localization is an essential problem as anomaly detection is. Because a rigorous localization requires a causal model of a target system, practically we often resort to a relaxed problem of anomaly interpretation, for which we are to obtain meaningful attribution of anomaly scores to input features. In this paper, we investigate the use of the Shapley value for anomaly interpretation. We focus on the semi-supervised anomaly detection and newly propose a characteristic function, on which the Shapley value is computed, specifically for anomaly scores. The idea of the proposed method is approximating the absence of some features by minimizing an anomaly score with regard to them. We examine the performance of the proposed method as well as other general approaches to computing the Shapley value in interpreting anomaly scores. We show the results of experiments on multiple datasets and anomaly detection methods, which indicate the usefulness of the Shapley-based anomaly interpretation toward anomaly localization.
研究动机与目标
- 为解决在缺乏完整因果模型的情况下解释异常分数的挑战。
- 开发一种有意义的异常分数归因方法,以支持异常定位。
- 提出一种专为异常分数设计的特征函数,以实现准确的Shapley值计算。
提出的方法
- 设计了一种新的特征函数,专门用于计算异常分数的Shapley值,通过最小化与这些特征相关的异常分数来建模特征缺失。
- 该方法采用基于最小化的近似方法,模拟输入特征的缺失,从而实现稳定且有意义的Shapley值估计。
- 在特征子集上计算Shapley值,以将异常分数归因于各个特征。
- 在多个异常检测模型和数据集上评估该方法,以评估其泛化能力和性能。
- 将该方法与通用的Shapley值计算技术进行比较,以验证其在异常解释中的有效性。
实验结果
研究问题
- RQ1在半监督异常检测中,Shapley值能否为异常分数提供有意义的特征归因?
- RQ2所提出的特征函数与通用方法相比,在捕捉相关特征贡献方面表现如何?
- RQ3基于Shapley值的解释是否能在多种数据集和模型上提升异常的可解释性和定位能力?
主要发现
- 所提出的基于Shapley值的方法在异常解释性能上优于通用的Shapley值计算方法。
- 通过Shapley值进行特征归因能有效突出对异常分数贡献最大的特征。
- 该方法在多个数据集和异常检测模型上均表现出良好的泛化能力。
- 基于最小化的特征缺失近似方法能够生成稳定且可解释的Shapley值。
- 实证结果表明,基于Shapley值的解释能够支持更精确的异常定位。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。