[论文解读] Do humans and Convolutional Neural Networks attend to similar areas during scene classification: Effects of task and image type
本研究探讨了在场景分类任务中,人类与卷积神经网络(CNN)的注意力机制如何随任务类型和图像内容而变化。通过在物体、室内场景和景观图像上进行眼动追踪与手动选择,研究发现人类任务意图显著影响其注意力与CNN注意力的相似性——尤其在物体图像中,手动选择与CNN的注意力最为一致;而在景观图像中,所有任务下的注意力相似性均较低。
Deep Learning models like Convolutional Neural Networks (CNN) are powerful image classifiers, but what factors determine whether they attend to similar image areas as humans do? While previous studies have focused on technological factors, little is known about the role of factors that affect human attention. In the present study, we investigated how the tasks used to elicit human attention maps interact with image characteristics in modulating the similarity between humans and CNN. We varied the intentionality of human tasks, ranging from spontaneous gaze during categorization over intentional gaze-pointing up to manual area selection. Moreover, we varied the type of image to be categorized, using either singular, salient objects, indoor scenes consisting of object arrangements, or landscapes without distinct objects defining the category. The human attention maps generated in this way were compared to the CNN attention maps revealed by explainable artificial intelligence (Grad-CAM). The influence of human tasks strongly depended on image type: For objects, human manual selection produced maps that were most similar to CNN, while the specific eye movement task has little impact. For indoor scenes, spontaneous gaze produced the least similarity, while for landscapes, similarity was equally low across all human tasks. To better understand these results, we also compared the different human attention maps to each other. Our results highlight the importance of taking human factors into account when comparing the attention of humans and CNN.
研究动机与目标
- 检验不同人类注意力任务(自发注视、注视指向、手动选择)如何影响人类注意力图与CNN之间的相似性。
- 探究图像类型(物体、室内场景、景观)如何调节人类与CNN注意力对齐的程度。
- 评估任务意图与图像内容是否共同塑造人类与模型注意力机制之间的相似性。
- 比较不同任务下的人类注意力图,以理解人类视觉注意力中的任务特异性偏差。
- 提供实证证据,说明在何种条件下人类与CNN注意力在场景分类过程中最为相似。
提出的方法
- 通过三种不同任务收集人类注意力图:在图像分类过程中进行自发眼动追踪、有意识的注视指向,以及手动区域选择。
- 使用Grad-CAM生成同一图像分类任务中预训练CNN的可解释注意力图。
- 通过相似性度量(如相关系数或交并比)比较人类与CNN注意力图在不同图像类型下的相似性。
- 系统性地改变图像内容:单个显著物体、具有物体布局的室内场景,以及无物体的景观。
- 通过多变量比较分析任务类型与图像类型对注意力相似性的影响。
- 评估不同任务下人类注意力图的一致性,以衡量任务的可靠性与偏差。
实验结果
研究问题
- RQ1人类注意力任务类型(自发注视、注视指向、手动选择)如何影响人类与CNN注意力图之间的相似性?
- RQ2图像类型(物体、室内场景、景观)如何影响人类与CNN注意力对齐的程度?
- RQ3在不同图像类型与任务模式下,人类-CNN注意力相似性是否存在一致的模式?
- RQ4不同任务下的人类注意力图如何相互比较?这对注意力对齐研究有何启示?
- RQ5在何种条件下,人类视觉注意力在场景分类任务中最接近CNN注意力?
主要发现
- 对于具有单一显著物体的图像,手动区域选择产生的人类注意力图与CNN Grad-CAM图最为相似,表明任务特异性对齐。
- 在室内场景中,自发注视与CNN注意力的相似性最低,表明自然眼动模式在复杂场景中与模型注意力存在显著差异。
- 对于缺乏显著物体的景观图像,所有人类任务下的注意力相似性均较低,表明物体缺失会降低对齐程度,且与任务类型无关。
- 人类任务的意图性具有强烈且依赖图像类型的影响:仅在物体和室内场景类别中显著影响相似性,而在景观中无显著影响。
- 人类注意力图在不同任务间存在显著差异,其中手动选择在定位相关图像区域方面表现出最高的一致性与特异性。
- 结果表明,人类因素(如任务设计与图像内容)在决定人类与CNN注意力对齐程度方面起着关键作用。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。