Skip to main content
QUICK REVIEW

[Paper Review] Do humans and Convolutional Neural Networks attend to similar areas during scene classification: Effects of task and image type

Romy Müller, Marcel Dürschmidt|arXiv (Cornell University)|Jul 25, 2023
Face Recognition and PerceptionNeuroscience70 references3 citations
TL;DR

This study investigates how human and Convolutional Neural Network (CNN) attention during scene classification vary with task type and image content. Using eye-tracking and manual selection across object, indoor scene, and landscape images, it finds that human task intentionality significantly affects attention similarity to CNNs—especially for objects, where manual selection aligns best with CNNs, while for landscapes, similarity remains low across all tasks.

ABSTRACT

Deep Learning models like Convolutional Neural Networks (CNN) are powerful image classifiers, but what factors determine whether they attend to similar image areas as humans do? While previous studies have focused on technological factors, little is known about the role of factors that affect human attention. In the present study, we investigated how the tasks used to elicit human attention maps interact with image characteristics in modulating the similarity between humans and CNN. We varied the intentionality of human tasks, ranging from spontaneous gaze during categorization over intentional gaze-pointing up to manual area selection. Moreover, we varied the type of image to be categorized, using either singular, salient objects, indoor scenes consisting of object arrangements, or landscapes without distinct objects defining the category. The human attention maps generated in this way were compared to the CNN attention maps revealed by explainable artificial intelligence (Grad-CAM). The influence of human tasks strongly depended on image type: For objects, human manual selection produced maps that were most similar to CNN, while the specific eye movement task has little impact. For indoor scenes, spontaneous gaze produced the least similarity, while for landscapes, similarity was equally low across all human tasks. To better understand these results, we also compared the different human attention maps to each other. Our results highlight the importance of taking human factors into account when comparing the attention of humans and CNN.

Motivation & Objective

  • To examine how different human attention tasks (spontaneous gaze, gaze-pointing, manual selection) influence attention map similarity with CNNs.
  • To investigate how image type (objects, indoor scenes, landscapes) modulates the degree of attention alignment between humans and CNNs.
  • To assess whether task intentionality and image content jointly shape the similarity between human and model attention mechanisms.
  • To compare human attention maps across different tasks to understand task-specific biases in human visual attention.
  • To provide empirical evidence on the conditions under which human and CNN attention are most similar during scene classification.

Proposed method

  • Collected human attention maps using three distinct tasks: spontaneous eye gaze during image categorization, intentional gaze-pointing, and manual area selection.
  • Used Grad-CAM to generate explainable attention maps from a pre-trained CNN for the same image classification task.
  • Compared human and CNN attention maps using similarity metrics (e.g., correlation or intersection over union) across different image types.
  • Systematically varied image content: singular salient objects, indoor scenes with object arrangements, and object-free landscapes.
  • Analyzed the influence of task type and image type on attention similarity through multivariate comparisons.
  • Evaluated inter-human attention map consistency across tasks to assess task reliability and bias.

Experimental results

Research questions

  • RQ1How does the type of human attention task (spontaneous gaze, gaze-pointing, manual selection) affect the similarity between human and CNN attention maps?
  • RQ2How does image type (objects, indoor scenes, landscapes) influence the degree of attention alignment between humans and CNNs?
  • RQ3Is there a consistent pattern in human-CNN attention similarity across different image types and task modalities?
  • RQ4How do human attention maps compare to one another across different tasks, and what does this imply for attention alignment studies?
  • RQ5Under what conditions is human visual attention most similar to CNN attention in scene classification tasks?

Key findings

  • For images with singular, salient objects, manual area selection produced human attention maps most similar to CNN Grad-CAM maps, indicating task-specific alignment.
  • In indoor scenes, spontaneous gaze resulted in the least similarity to CNN attention, suggesting that natural gaze patterns diverge from model attention in complex scenes.
  • For landscapes lacking distinct objects, attention similarity remained low across all human tasks, indicating that object absence reduces alignment regardless of task type.
  • The intentionality of the human task had a strong, image-type-dependent effect: it significantly influenced similarity only in object and indoor scene categories, not in landscapes.
  • Human attention maps varied substantially across tasks, with manual selection showing the highest consistency and specificity in targeting relevant image regions.
  • The results demonstrate that human factors such as task design and image content are critical in determining the degree of attention alignment between humans and CNNs.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.