Skip to main content
QUICK REVIEW

[论文解读] Exemplary Natural Images Explain CNN Activations Better than State-of-the-Art Feature Visualization

Judy Borowski, R. Zimmermann|arXiv (Cornell University)|Oct 23, 2020
Explainable Artificial Intelligence (XAI)被引用 13
一句话总结

本研究通过将Olah等人(2017年)提出的合成特征可视化结果与能强烈激活相同特征的自然样本图像进行对比,评估了卷积神经网络(CNN)激活解释的可解释性。通过一项受控的心理物理学实验,发现当人类参与者以自然图像作为参考时,其在识别强激活图像方面的准确率显著更高(92±2%),优于广泛用于可解释性研究的合成可视化结果(82±4%)。

ABSTRACT

Feature visualizations such as synthetic maximally activating images are a widely used explanation method to better understand the information processing of convolutional neural networks (CNNs). At the same time, there are concerns that these visualizations might not accurately represent CNNs' inner workings. Here, we measure how much extremely activating images help humans to predict CNN activations. Using a well-controlled psychophysical paradigm, we compare the informativeness of synthetic images by Olah et al. (2017) with a simple baseline visualization, namely exemplary natural images that also strongly activate a specific feature map. Given either synthetic or natural reference images, human participants choose which of two query images leads to strong positive activation. The experiments are designed to maximize participants' performance, and are the first to probe intermediate instead of final layer representations. We find that synthetic images indeed provide helpful information about feature map activations ($82\\pm4\\%$ accuracy; chance would be $50\\%$). However, natural images - originally intended as a baseline - outperform synthetic images by a wide margin ($92\\pm2\\%$). Additionally, participants are faster and more confident for natural images, whereas subjective impressions about the interpretability of the feature visualizations are mixed. The higher informativeness of natural images holds across most layers, for both expert and lay participants as well as for hand- and randomly-picked feature visualizations. Even if only a single reference image is given, synthetic images provide less information than natural images ($65\\pm5\\%$ vs. $73\\pm4\\%$). In summary, synthetic images from a popular feature visualization method are significantly less informative for assessing CNN activations than natural images. We argue that visualization methods should improve over this baseline.

研究动机与目标

  • 评估合成特征可视化与自然样本图像在解释CNN特征图激活方面的信息量。
  • 探究Olah等人(2017年)提出的广泛使用的合成可视化方法是否真正有效于人类对CNN行为的理解。
  • 评估原本用作简单基线的自然图像是否对专家和普通用户均提供更优的可解释性。
  • 衡量可视化类型对人类在激活预测任务中的判断准确率、反应时间及信心评分的影响。
  • 确定自然图像的优势是否在不同网络层中持续存在,且对人工挑选与随机选择的特征图均成立。

提出的方法

  • 开展一项受控的心理物理学实验,让人类参与者判断两个查询图像中哪一个在特定CNN特征图上引发更强的激活。
  • 使用两种类型的参考图像:来自Olah等人(2017年)的合成最大激活图像,以及同样强烈激活同一特征图的真实自然图像(ImageNet数据集)。
  • 设计任务以最大程度提升表现,通过最小化模糊性并确保查询图像之间的激活差异清晰可辨。
  • 收集专家和普通参与者在判断准确率、反应时间及信心评分方面的数据。
  • 在神经网络的多个层级上评估表现,并针对人工挑选与随机选择的特征图进行分析。
  • 通过仅提供单一参考图像的消融研究,评估信息增益的最小限度。

实验结果

研究问题

  • RQ1Olah等人(2017年)提出的合成特征可视化是否比自然样本图像为CNN激活提供更具信息量的解释?
  • RQ2当使用合成与自然参考图像时,人类参与者在预测CNN激活方面的表现有何差异?
  • RQ3自然图像的优势是否在不同网络层中持续存在,且对专家与普通用户均成立?
  • RQ4使用合成与自然可视化时,反应时间与信心评分有何不同?
  • RQ5当仅提供单一参考图像时,自然图像带来的性能提升是否依然显著?

主要发现

  • 当以自然样本图像作为参考时,参与者在识别强激活图像方面的准确率达到92±2%,显著高于使用合成可视化时的82±4%。
  • 该性能差距在所有网络层中均持续存在,包括中间层,表明自然图像的优势并非局限于深层或浅层特征。
  • 参与者使用自然图像作为参考时反应更快、信心更高,表明其认知处理与可解释性更优。
  • 即使仅提供单一参考图像,自然图像仍提供更多信息(73±4%准确率),优于合成图像(65±5%准确率),证明其作为可解释性工具的稳健性。
  • 自然图像的优势在人工挑选与随机选择的特征图中均成立,表明其相对于合成可视化具有普遍性优势。
  • 主观评估显示,对合成可视化可解释性的评价褒贬不一,而自然图像则始终被认为更具意义且更直观。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。