[论文解读] Exemplary Natural Images Explain CNN Activations Better than Feature Visualizations
本研究通过对比合成特征可视化与自然样本图像在解释CNN激活方面的表现,开展了一项受控的心理物理学实验。结果表明,自然图像在帮助人类预测特征图响应方面优于合成可视化方法(准确率分别为92%与82%),说明自然图像更具信息量,应作为未来可视化方法的基线标准。
Feature visualizations such as synthetic maximally activating images are a widely used explanation method to better understand the information processing of convolutional neural networks (CNNs). At the same time, there are concerns that these visualizations might not accurately represent CNNs' inner workings. Here, we measure how much extremely activating images help humans to predict CNN activations. Using a well-controlled psychophysical paradigm, we compare the informativeness of synthetic images (Olah et al., 2017) with a simple baseline visualization, namely exemplary natural images that also strongly activate a specific feature map. Given either synthetic or natural reference images, human participants choose which of two query images leads to strong positive activation. The experiment is designed to maximize participants' performance, and is the first to probe intermediate instead of final layer representations. We find that synthetic images indeed provide helpful information about feature map activations (82% accuracy; chance would be 50%). However, natural images-originally intended to be a baseline-outperform synthetic images by a wide margin (92% accuracy). Additionally, participants are faster and more confident for natural images, whereas subjective impressions about the interpretability of feature visualization are mixed. The higher informativeness of natural images holds across most layers, for both expert and lay participants as well as for hand- and randomly-picked feature visualizations. Even if only a single reference image is given, synthetic images provide less information than natural images (65% vs. 73%). In summary, popular synthetic images from feature visualizations are significantly less informative for assessing CNN activations than natural images. We argue that future visualization methods should improve over this simple baseline.
研究动机与目标
- 评估合成特征可视化与自然样本图像在解释CNN特征图激活方面的有效性。
- 探究原本用作基线的自然图像是否比合成可视化为理解CNN行为提供更具信息量的线索。
- 比较人类在使用合成与自然参考图像时,对不同网络层及不同参与者群体的CNN激活预测表现。
- 评估参考图像类型对人类反应速度、信心水平及主观可解释性的影响。
提出的方法
- 开展了一项受控的心理物理学实验,参与者需判断两个查询图像中哪一个更强烈地激活了特定的CNN特征图。
- 同时使用合成的最激活图像(Olah et al., 2017)和真实自然图像作为参考刺激,这些自然图像同样强烈激活了相同的特征图。
- 通过二分类任务测量人类在预测特征图激活强度方面的表现,准确率为主要指标。
- 测试了网络的最终层与中间层,以评估结果在不同网络深度下的泛化能力。
- 收集了专家与非专家参与者的数据,以评估在不同用户专业水平下的泛化性。
- 通过单参考与双参考条件的对比,分离出图像类型的影响。
实验结果
研究问题
- RQ1合成特征可视化是否比自然样本图像更有效地帮助人类预测CNN特征图激活?
- RQ2人类在使用合成与自然参考图像时,预测CNN激活的表现有何差异?
- RQ3自然图像相较于合成可视化的优势是否在不同CNN层及不同参与者专业水平下均成立?
- RQ4使用自然图像与合成图像时,反应时间与信心水平有何不同?
- RQ5单张自然图像是否能比合成图像提供更多的预测信息,用于理解CNN激活模式?
主要发现
- 自然样本图像在预测CNN激活方面显著优于合成特征可视化,准确率分别达到92%与82%。
- 自然图像的优势在大多数CNN层中保持一致,包括中间层。
- 参与者使用自然图像作为参考时反应更快、信心更高,表明其可用性更优。
- 即使仅使用单张参考图像,自然图像的预测准确率(73%)也高于合成图像(65%)。
- 该优势在专家与非专家参与者中均成立,表明其具有广泛适用性。
- 对于合成可视化,主观可解释性评价结果不一;而自然图像始终被认为更具直观性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。