Skip to main content
QUICK REVIEW

[Paper Review] Exemplary Natural Images Explain CNN Activations Better than Feature Visualizations

Judy Borowski, R. Zimmermann|arXiv (Cornell University)|Oct 23, 2020
Explainable Artificial Intelligence (XAI)4 citations
TL;DR

This study compares synthetic feature visualizations with natural exemplar images in explaining CNN activations. Using a controlled psychophysical experiment, it finds that natural images outperform synthetic visualizations in helping humans predict feature map responses (92% vs. 82% accuracy), suggesting natural images are more informative and should serve as a baseline for future visualization methods.

ABSTRACT

Feature visualizations such as synthetic maximally activating images are a widely used explanation method to better understand the information processing of convolutional neural networks (CNNs). At the same time, there are concerns that these visualizations might not accurately represent CNNs' inner workings. Here, we measure how much extremely activating images help humans to predict CNN activations. Using a well-controlled psychophysical paradigm, we compare the informativeness of synthetic images (Olah et al., 2017) with a simple baseline visualization, namely exemplary natural images that also strongly activate a specific feature map. Given either synthetic or natural reference images, human participants choose which of two query images leads to strong positive activation. The experiment is designed to maximize participants' performance, and is the first to probe intermediate instead of final layer representations. We find that synthetic images indeed provide helpful information about feature map activations (82% accuracy; chance would be 50%). However, natural images-originally intended to be a baseline-outperform synthetic images by a wide margin (92% accuracy). Additionally, participants are faster and more confident for natural images, whereas subjective impressions about the interpretability of feature visualization are mixed. The higher informativeness of natural images holds across most layers, for both expert and lay participants as well as for hand- and randomly-picked feature visualizations. Even if only a single reference image is given, synthetic images provide less information than natural images (65% vs. 73%). In summary, popular synthetic images from feature visualizations are significantly less informative for assessing CNN activations than natural images. We argue that future visualization methods should improve over this simple baseline.

Motivation & Objective

  • To evaluate how effectively synthetic feature visualizations and natural exemplar images explain CNN feature map activations.
  • To investigate whether natural images—originally intended as a baseline—provide more informative cues than synthetic visualizations for understanding CNN behavior.
  • To compare human performance in predicting CNN activations using synthetic versus natural reference images across different network layers and participant groups.
  • To assess the impact of reference image type on human speed, confidence, and subjective interpretability.

Proposed method

  • Conducted a well-controlled psychophysical experiment where participants judged which of two query images activated a specific CNN feature map more strongly.
  • Used both synthetic maximally activating images (Olah et al., 2017) and real natural images that strongly activate the same feature map as reference stimuli.
  • Measured human performance in binary classification tasks to predict feature map activation strength, with accuracy as the primary metric.
  • Tested both final and intermediate CNN layers to assess generalizability across network depth.
  • Collected data from both expert and lay participants to evaluate generalizability across user expertise.
  • Compared performance with single-reference and dual-reference conditions to isolate the effect of image type.

Experimental results

Research questions

  • RQ1Do synthetic feature visualizations help humans predict CNN feature map activations more accurately than natural exemplar images?
  • RQ2How does the performance of human participants in predicting CNN activations vary between synthetic and natural reference images?
  • RQ3Does the advantage of natural images over synthetic visualizations hold across different CNN layers and participant expertise levels?
  • RQ4How do response time and confidence differ when using natural versus synthetic reference images?
  • RQ5Can a single natural image provide more predictive information than a synthetic image for CNN activation patterns?

Key findings

  • Natural exemplar images significantly outperformed synthetic feature visualizations in predicting CNN activations, achieving 92% accuracy compared to 82% for synthetic images.
  • The performance advantage of natural images was consistent across most CNN layers, including intermediate layers.
  • Participants were faster and more confident when using natural images as references, indicating improved usability.
  • Even with only a single reference image, natural images provided better predictive performance (73% accuracy) than synthetic images (65% accuracy).
  • The advantage of natural images held across both expert and lay participants, suggesting broad applicability.
  • Subjective impressions of interpretability were mixed for synthetic visualizations, while natural images were consistently perceived as more intuitive.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.