[Paper Review] Adversarially trained neural representations may already be as robust as corresponding biological neural representations
This study introduces a method to perform adversarial attacks directly on primate inferior temporal (IT) cortex neural activity, revealing that biological neural representations in primates exhibit adversarial sensitivity comparable to state-of-the-art robustly trained artificial neural networks. The key finding is that IT neurons are vulnerable to small $l_2$-norm perturbations (as low as $\epsilon = 2.5$), challenging the long-held assumption that biological vision is inherently more robust than artificial systems.
Visual systems of primates are the gold standard of robust perception. There is thus a general belief that mimicking the neural representations that underlie those systems will yield artificial visual systems that are adversarially robust. In this work, we develop a method for performing adversarial visual attacks directly on primate brain activity. We then leverage this method to demonstrate that the above-mentioned belief might not be well founded. Specifically, we report that the biological neurons that make up visual systems of primates exhibit susceptibility to adversarial perturbations that is comparable in magnitude to existing (robustly trained) artificial neural networks.
Motivation & Objective
- To test whether high-level biological neural representations in the primate visual system are truly more robust to adversarial perturbations than artificial neural networks.
- To overcome the limitation of prior neuroscience studies that relied on random or coarse image perturbations by developing a targeted, iterative attack method on neural recordings.
- To directly compare the adversarial sensitivity of individual IT neurons with units in robustly trained deep neural networks under controlled $l_2$-norm perturbation budgets.
- To investigate whether the perceived robustness of primate visual perception stems from population-level error correction rather than individual neuron resilience.
Proposed method
- Developed an iterative projected gradient descent (PGD) attack tailored for primate neural recordings, using the neural response as a surrogate loss to optimize adversarial perturbations.
- Applied a two-phase PGD strategy: first optimize with a larger $l_2$-ball ($2\epsilon$) to improve exploration, then refine with $\epsilon$ to achieve precise perturbations.
- Used simulated annealing with restarts, reducing step size by 10% when no progress was made, repeated up to four times to escape local minima.
- Leveraged mechanistic models of primate visual processing to generate image perturbations that maximize response shifts in individual IT neurons.
- Optimized for worst-case perturbations under $l_2$-norm constraints, ensuring adversarial sensitivity was measured as a lower bound on neuronal response change.
- Validated results by comparing adversarial sensitivity across IT neurons, standard ResNet50, and adversarially trained networks (AT-ResNet50, AT-WRN50-4) at $\epsilon = 3$.
Experimental results
Research questions
- RQ1Are primate IT neurons vulnerable to small $l_2$-norm adversarial perturbations similar in magnitude to those used in deep learning robustness research?
- RQ2How does the adversarial sensitivity of individual IT neurons compare quantitatively to that of units in state-of-the-art robustly trained artificial neural networks?
- RQ3Is the robustness of primate visual perception due to individual neuron resilience or population-level error correction mechanisms?
- RQ4Can adversarial attacks be effectively applied to biological neural recordings using gradient-based optimization, given the lack of full model access?
- RQ5Do adversarial examples in IT neurons exist densely in image space, such that any image can be perturbed to alter category selectivity?
Key findings
- Primate IT neurons exhibit adversarial sensitivity to $l_2$-norm perturbations as small as $\epsilon = 2.5$, which are nearly imperceptible to humans.
- The average adversarial sensitivity of IT neurons at $\epsilon = 3$ is comparable in magnitude to that of adversarially trained DNNs (AT-ResNet50, AT-WRN50-4), challenging the assumption of biological superiority in robustness.
- Adversarial perturbations can completely alter the category selectivity of individual IT neurons, indicating that their response is highly sensitive to small, targeted pixel changes.
- The study demonstrates that adversarial examples exist densely in image space around any given image, suggesting widespread vulnerability across the visual representation.
- Robustly trained artificial neural networks have already achieved adversarial robustness levels matching or exceeding those of biological neural representations at the individual unit level.
- The results imply that the robustness of primate visual behavior may not stem from individual neuron resilience but from downstream population-level decoding or error-correction mechanisms.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.