Skip to main content
QUICK REVIEW

[Paper Review] How Deep is the Feature Analysis underlying Rapid Visual Categorization?

Sven Eberhardt, Jonah Cader|arXiv (Cornell University)|Jun 3, 2016
Face Recognition and Perception16 references17 citations
TL;DR

This study investigates the depth of visual feature processing in rapid human categorization by comparing deep neural network activations to human behavioral responses on an animal vs. non-animal categorization task. It finds that human decisions correlate best with intermediate-level features in deep networks—suggesting humans use moderately complex representations—while deeper network layers exceed human performance but diverge in their decision patterns.

ABSTRACT

Rapid categorization paradigms have a long history in experimental psychology: Characterized by short presentation times and speedy behavioral responses, these tasks highlight the efficiency with which our visual system processes natural object categories. Previous studies have shown that feed-forward hierarchical models of the visual cortex provide a good fit to human visual decisions. At the same time, recent work in computer vision has demonstrated significant gains in object recognition accuracy with increasingly deep hierarchical architectures. But it is unclear how well these models account for human visual decisions and what they may reveal about the underlying brain processes. We have conducted a large-scale psychophysics study to assess the correlation between computational models and human participants on a rapid animal vs. non-animal categorization task. We considered visual representations of varying complexity by analyzing the output of different stages of processing in three state-of-the-art deep networks. We found that recognition accuracy increases with higher stages of visual processing (higher level stages indeed outperforming human participants on the same task) but that human decisions agree best with predictions from intermediate stages. Overall, these results suggest that human participants may rely on visual features of intermediate complexity and that the complexity of visual representations afforded by modern deep network models may exceed those used by human participants during rapid categorization.

Motivation & Objective

  • To determine the depth of visual feature analysis underlying rapid human visual categorization.
  • To assess whether modern deep neural networks, which outperform humans in object recognition, accurately model human decision-making in rapid categorization.
  • To identify at which stage of visual processing human decisions align best with computational model predictions.
  • To investigate whether increased network depth leads to better modeling of human behavior or diverges from human strategies.

Proposed method

  • Conducted a large-scale psychophysics experiment with 281 participants on Amazon Mechanical Turk using 2,100 balanced animal and non-animal images.
  • Used three state-of-the-art deep networks—AlexNet, VGG16, and VGG19—trained on ImageNet for feature extraction.
  • Extracted and analyzed activation responses from individual layers across all networks to assess recognition accuracy and human-model agreement.
  • Measured correlation between model predictions and human responses at each layer, focusing on intermediate vs. deep processing stages.
  • Performed a control experiment with varying response times (500ms, 1000ms, 2000ms) to test the impact of time on human-model agreement.
  • Applied statistical analysis (one-way ANOVA with Tukey’s HSD) to evaluate differences in categorization accuracy across time conditions.

Experimental results

Research questions

  • RQ1At which stage of visual processing do human decisions on rapid categorization tasks most closely align with deep neural network predictions?
  • RQ2Does the increasing depth of modern deep networks improve their ability to model human rapid categorization behavior or do they diverge?
  • RQ3How does response time affect the correlation between human decisions and model predictions across different network layers?
  • RQ4Are there specific image types where deep networks outperform or underperform humans, and what visual features drive these differences?

Key findings

  • Recognition accuracy of deep network layers increased monotonically with depth across all models, with the deepest layers outperforming human participants on the same task.
  • Human decisions showed maximum correlation with model predictions at intermediate layers—specifically around layer conv5_2 in VGG16—after which agreement decreased.
  • The peak correlation between human and model decisions occurred at approximately 10 layers of processing, consistent with the number of stages in the ventral visual stream.
  • Deeper layers showed higher accuracy but lower agreement with human responses, particularly on elongated animals (e.g., snakes), camouflaged animals, and atypical object contexts.
  • Humans outperformed deep networks on iconic, direct-view illustrations (e.g., a cat facing the camera), indicating reliance on high-level, familiar visual configurations.
  • Extending response time from 500ms to 1000ms significantly improved human accuracy (from 74% to 84%), but no further gain occurred at 2000ms, and the pattern of human-model agreement remained stable across time conditions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.