Skip to main content
QUICK REVIEW

[论文解读] Passive Attention in Artificial Neural Networks Predicts Human Visual Selectivity

Thomas A. Langlois, He Zhao|arXiv (Cornell University)|Jul 14, 2021
Visual Attention and Saliency Detection参考文献 43被引用 4
一句话总结

本研究表明,人工神经网络(ANNs)中的被动注意力机制,特别是简单架构中的引导反向传播(guided backpropagation),能够预测六项行为任务中的人类视觉选择性。该研究展示了ANN注意力图与人类注意力之间存在强烈相关性,通过因果识别实验得到验证,其中基于ANN和人类的掩码均提升了分类性能。

ABSTRACT

Developments in machine learning interpretability techniques over the past decade have provided new tools to observe the image regions that are most informative for classification and localization in artificial neural networks (ANNs). Are the same regions similarly informative to human observers? Using data from 79 new experiments and 7,810 participants, we show that passive attention techniques reveal a significant overlap with human visual selectivity estimates derived from 6 distinct behavioral tasks including visual discrimination, spatial localization, recognizability, free-viewing, cued-object search, and saliency search fixations. We find that input visualizations derived from relatively simple ANN architectures probed using guided backpropagation methods are the best predictors of a shared component in the joint variability of the human measures. We validate these correlational results with causal manipulations using recognition experiments. We show that images masked with ANN attention maps were easier for humans to classify than control masks in a speeded recognition experiment. Similarly, we find that recognition performance in the same ANN models was likewise influenced by masking input images using human visual selectivity maps. This work contributes a new approach to evaluating the biological and psychological validity of leading ANNs as models of human vision: by examining their similarities and differences in terms of their visual selectivity to the information contained in images.

研究动机与目标

  • 评估ANN中的被动注意力机制是否能预测多种行为任务中的人类视觉选择性。
  • 比较不同ANN可解释性技术对人类视觉注意力多种行为测量指标的预测能力。
  • 通过因果识别实验验证ANN注意力与人类视觉选择性之间的相关性。
  • 评估不同行为测量指标(如辨别、定位、可识别性)是否捕捉到不同的视觉信息,并在与ANN注意力的一致性上是否存在差异。

提出的方法

  • 在79项新实验中收集了7,910名参与者的实验数据,涵盖六项行为任务:视觉辨别、空间定位、可识别性、自由观察、提示目标搜索和显著性搜索注视点。
  • 对10种不同ANN架构应用被动注意力技术——特别是引导反向传播(SGBP)——以提取对分类最具影响力的视觉区域。
  • 使用高斯平滑处理,采用学习得到的参数σ将ANN注意力图与人类行为图对齐,通过交叉验证在训练集上优化σ。
  • 采用分半交叉验证以验证相关性结果的稳健性,即在训练集上拟合σ,并在保留集上进行测试。
  • 开展快速识别实验以检验因果影响:使用ANN注意力图和人类视觉选择性图对图像进行掩码处理,随后测量人类和模型的分类性能。
  • 在所有条件下计算ANN注意力图与人类行为图(如人类PC、补丁评分、辨别准确率)之间的峰值相关性。

实验结果

研究问题

  • RQ1ANN中的被动注意力图是否能预测多种行为任务中的人类视觉选择性?
  • RQ2哪种ANN可解释性方法和架构最能预测人类视觉注意力的共享变异?
  • RQ3通过识别性能是否能证明ANN注意力与人类注意力之间的对应关系具有因果意义?
  • RQ4不同的人类行为测量指标(如辨别与定位)在与ANN注意力图的一致性上是否存在差异?
  • RQ5相关性结果在不同数据划分和平滑参数下是否具有稳健性?

主要发现

  • 在简单ANN(如AlexNet)中,引导反向传播产生的峰值相关性最高(r ≈ 0.71),优于其他方法和架构。
  • ANN注意力图在所有六项行为任务中均显著预测了人类视觉选择性,其中与人类PC、补丁评分和辨别准确率图的相关性最强。
  • 在快速识别实验中,使用ANN注意力图掩码的图像比使用控制掩码(如随机或显著性掩码)的人类分类更快且更准确。
  • 同一ANN的分类性能也显著受到人类视觉选择性图的影响,证实了双向一致性。
  • 分半交叉验证结果表明结果稳健:测试集注意力图的峰值相关性(r ≈ 0.71)与训练集结果(r ≈ 0.71)高度一致,100次随机划分中方差极小。
  • 存在显著交互作用(F(6,176)=6.54, p < 0.001),表明与ANN注意力相关性更高的行为图(如PC、补丁评分)在正确与错误掩码条件下的倒序排名性能差异更大。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。