Skip to main content
QUICK REVIEW

[论文解读] Quantifying Statistical Significance of Neural Network Representation-Driven Hypotheses by Selective Inference

Vo Nguyen Le Duy, Shogo Iwazaki|arXiv (Cornell University)|May 4, 2021
Explainable Artificial Intelligence (XAI)参考文献 30被引用 13
一句话总结

该论文提出了一种选择性推断框架,用于量化基于深度神经网络(DNN)表征推导出的假设的统计显著性,采用基于同伦的算法计算精确的条件抽样分布。该方法控制了假阳性率,并在合成数据集和真实世界数据集上展现出强大的计算效率和实际性能。

ABSTRACT

In the past few years, various approaches have been developed to explain and interpret deep neural network (DNN) representations, but it has been pointed out that these representations are sometimes unstable and not reproducible. In this paper, we interpret these representations as hypotheses driven by DNN (called DNN-driven hypotheses) and propose a method to quantify the reliability of these hypotheses in statistical hypothesis testing framework. To this end, we introduce Selective Inference (SI) framework, which has received much attention in the past few years as a new statistical inference framework for data-driven hypotheses. The basic idea of SI is to make conditional inferences on the selected hypotheses under the condition that they are selected. In order to use SI framework for DNN representations, we develop a new SI algorithm based on homotopy method which enables us to derive the exact (non-asymptotic) conditional sampling distribution of the DNN-driven hypotheses. We conduct experiments on both synthetic and real-world datasets, through which we offer evidence that our proposed method can successfully control the false positive rate, has decent performance in terms of computational efficiency, and provides good results in practical applications.

研究动机与目标

  • 解决深度神经网络(DNN)表征解释中的不稳定性和缺乏可复现性问题。
  • 将DNN驱动的解释重新定义为需要统计验证的数据驱动假设。
  • 开发一个可靠的统计框架,以评估这些假设的显著性,超越启发式或渐近方法。
  • 实现对DNN驱动假设的精确、非渐近推断,以提升可解释性和可信度。

提出的方法

  • 采用选择性推断(SI)框架,对基于DNN表征选择的假设进行条件推断。
  • 开发一种新颖的基于同伦的算法,以计算所选DNN驱动假设的精确(非渐近)条件抽样分布。
  • 利用同伦方法中的路径跟踪技术,高效追踪解路径,并在选择后推导出有效的p值。
  • 将同伦算法集成到DNN表征流程中,以实现对特征重要性和表征驱动假设的统计检验。
  • 通过条件化于选择事件来确保控制第一类错误率,避免事后选择推断的陷阱。

实验结果

研究问题

  • RQ1选择性推断能否有效应用于量化基于DNN表征推导出的假设的统计显著性?
  • RQ2所提出的基于同伦的SI方法能否为DNN驱动的假设提供精确的、非渐近的推断?
  • RQ3与传统方法相比,该方法在实践中对假阳性率的控制效果如何?
  • RQ4该方法在真实世界和合成数据集上的计算效率和可扩展性如何?

主要发现

  • 所提出的方法在合成数据集和真实世界数据集上均成功控制了假阳性率,验证了其统计可靠性。
  • 基于同伦的算法实现了无需依赖渐近近似的精确条件推断,相较于标准方法提高了准确性。
  • 该方法展现出良好的计算效率,使其在可解释性流程中具备实际部署的可行性。
  • 实证结果表明,该方法在多种数据环境下均能产生有意义且可复现的DNN表征解释。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。