Skip to main content
QUICK REVIEW

[Paper Review] Passive Attention in Artificial Neural Networks Predicts Human Visual Selectivity

Thomas A. Langlois, He Zhao|arXiv (Cornell University)|Jul 14, 2021
Visual Attention and Saliency Detection43 references4 citations
TL;DR

This study demonstrates that passive attention mechanisms in artificial neural networks (ANNs), particularly guided backpropagation in simple architectures, predict human visual selectivity across six behavioral tasks. It shows strong correlation between ANN attention maps and human attention, validated through causal recognition experiments where ANN- and human-based masks improved classification performance.

ABSTRACT

Developments in machine learning interpretability techniques over the past decade have provided new tools to observe the image regions that are most informative for classification and localization in artificial neural networks (ANNs). Are the same regions similarly informative to human observers? Using data from 79 new experiments and 7,810 participants, we show that passive attention techniques reveal a significant overlap with human visual selectivity estimates derived from 6 distinct behavioral tasks including visual discrimination, spatial localization, recognizability, free-viewing, cued-object search, and saliency search fixations. We find that input visualizations derived from relatively simple ANN architectures probed using guided backpropagation methods are the best predictors of a shared component in the joint variability of the human measures. We validate these correlational results with causal manipulations using recognition experiments. We show that images masked with ANN attention maps were easier for humans to classify than control masks in a speeded recognition experiment. Similarly, we find that recognition performance in the same ANN models was likewise influenced by masking input images using human visual selectivity maps. This work contributes a new approach to evaluating the biological and psychological validity of leading ANNs as models of human vision: by examining their similarities and differences in terms of their visual selectivity to the information contained in images.

Motivation & Objective

  • To evaluate whether passive attention mechanisms in ANNs predict human visual selectivity across diverse behavioral tasks.
  • To compare the predictive power of various ANN interpretability techniques against multiple human behavioral measures of visual attention.
  • To validate the correlation between ANN attention and human visual selectivity through causal recognition experiments.
  • To assess whether different behavioral measures (e.g., discrimination, localization, recognizability) capture distinct visual information and vary in their alignment with ANN attention.

Proposed method

  • Collected data from 7,910 participants across 79 new experiments using six behavioral tasks: visual discrimination, spatial localization, recognizability, free-viewing, cued-object search, and saliency search fixations.
  • Applied passive attention techniques—specifically guided backpropagation (SGBP)—to extract visual regions most influential for classification in 10 different ANN architectures.
  • Used Gaussian smoothing with a learned parameter σ to align ANN attention maps with human behavioral maps, optimizing σ on training splits via cross-validation.
  • Performed split-half cross-validation to validate robustness of correlation results, fitting σ on training sets and testing on held-out sets.
  • Conducted speeded recognition experiments to test causal influence: masked images using ANN attention maps and human visual selectivity maps, then measured human and model classification performance.
  • Calculated peak correlations between ANN attention maps and human behavioral maps (e.g., human PC, patch ratings, discrimination accuracy) across all conditions.

Experimental results

Research questions

  • RQ1Do passive attention maps in ANNs predict human visual selectivity across multiple behavioral tasks?
  • RQ2Which ANN interpretability method and architecture best predict shared variability in human visual attention?
  • RQ3Is the correspondence between ANN attention and human attention causally meaningful, as shown by recognition performance?
  • RQ4Do different human behavioral measures (e.g., discrimination vs. localization) vary in their alignment with ANN attention maps?
  • RQ5How robust are the correlation results across different data splits and smoothing parameters?

Key findings

  • Guided backpropagation in simple ANNs (e.g., AlexNet) produced the highest peak correlations (r ≈ 0.71) with human perceptual maps, outperforming other methods and architectures.
  • ANN attention maps significantly predicted human visual selectivity across all six behavioral tasks, with the strongest correlations observed for human PC, patch ratings, and discrimination accuracy maps.
  • In speeded recognition experiments, images masked with ANN attention maps were classified faster and more accurately by humans than those masked with control (e.g., random or saliency-based) masks.
  • Recognition performance in the same ANNs was also significantly influenced by human visual selectivity maps, confirming bidirectional alignment.
  • Split-half cross-validation confirmed robustness: peak correlations for test set maps (r ≈ 0.71) closely matched training set results (r ≈ 0.71), with minimal variance across 100 random splits.
  • A significant interaction (F(6,176)=6.54, p < 0.001) showed that behavioral maps with higher correlation to ANN attention (e.g., PC, patch ratings) produced larger differences in inverse-rank performance between correct and incorrect masking conditions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.