[Paper Review] Adversarial Examples that Fool both Computer Vision and Time-Limited Humans
Adversarial perturbations that transfer across CNNs can bias time-limited human judgments and increase error rates, revealing shared failure modes between machines and human vision.
Machine learning models are vulnerable to adversarial examples: small changes to images can cause computer vision models to make mistakes such as identifying a school bus as an ostrich. However, it is still an open question whether humans are prone to similar mistakes. Here, we address this question by leveraging recent techniques that transfer adversarial examples from computer vision models with known parameters and architecture to other models with unknown parameters and architecture, and by matching the initial processing of the human visual system. We find that adversarial examples that strongly transfer across computer vision models influence the classifications made by time-limited human observers.
Motivation & Objective
- Investigate whether adversarial examples that fool computer vision models also affect human perception under time constraints.
- Bridge machine learning and neuroscience by aligning early human visual processing with CNN inputs.
- Measure transferability of adversarial perturbations from ensemble CNNs to time-limited human observers.
Proposed method
- Construct an ensemble of 10 CNN models (Inception and ResNet variants) with a retinal preprocessing layer to mimic human early vision.
- Generate targeted adversarial perturbations with bounded L-infinity norm to cause misclassification across the model ensemble.
- Use a black-box adversarial attack approach that does not require access to model architectures or parameters.
- Present images briefly with masks to time-limited human subjects to mimic feedforward processing and limit top-down influences.
- Evaluate human decisions in a two-alternative forced-choice task across multiple image groups (pets, vegetables, hazards).
- Compare adversarial effects against control conditions (image and flip) and a false condition to isolate perceptual influence.
Experimental results
Research questions
- RQ1Do adversarial examples that transfer between CNNs also bias time-limited human perception?
- RQ2How does retinal-like preprocessing affect transfer of adversarial perturbations to humans?
- RQ3What is the impact of adversarial perturbations on human accuracy and decision time under brief presentation?
- RQ4Are adversarial perturbations capable of forcing incorrect choices when the true class is an available option?
Key findings
- Adversarial perturbations transferred to time-limited humans bias choices toward targeted incorrect classes.
- Adversarial images reduced human accuracy relative to clean images, with perturbations stronger than a vertical flip control.
- Response times increased in the adversarial condition, and quicker decisions showed stronger bias toward targeted classes.
- Transfer success to humans varied by image group, with hazard images showing stronger bias than pets, which were stronger than vegetables.
- Even when the correct class was available, adversarial perturbations increased error rates beyond the baseline image condition.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.