[Paper Review] Performance-optimized deep neural networks are evolving into worse models of inferotemporal visual cortex
The paper shows that as DNNs improve on ImageNet, they become worse at predicting IT neural responses; training with the neural harmonizer aligns representations with humans and restores neural predictivity.
One of the most impactful findings in computational neuroscience over the past decade is that the object recognition accuracy of deep neural networks (DNNs) correlates with their ability to predict neural responses to natural images in the inferotemporal (IT) cortex. This discovery supported the long-held theory that object recognition is a core objective of the visual cortex, and suggested that more accurate DNNs would serve as better models of IT neuron responses to images. Since then, deep learning has undergone a revolution of scale: billion parameter-scale DNNs trained on billions of images are rivaling or outperforming humans at visual tasks including object recognition. Have today's DNNs become more accurate at predicting IT neuron responses to images as they have grown more accurate at object recognition? Surprisingly, across three independent experiments, we find this is not the case. DNNs have become progressively worse models of IT as their accuracy has increased on ImageNet. To understand why DNNs experience this trade-off and evaluate if they are still an appropriate paradigm for modeling the visual system, we turn to recordings of IT that capture spatially resolved maps of neuronal activity elicited by natural images. These neuronal activity maps reveal that DNNs trained on ImageNet learn to rely on different visual features than those encoded by IT and that this problem worsens as their accuracy increases. We successfully resolved this issue with the neural harmonizer, a plug-and-play training routine for DNNs that aligns their learned representations with humans. Our results suggest that harmonized DNNs break the trade-off between ImageNet accuracy and neural prediction accuracy that assails current DNNs and offer a path to more accurate models of biological vision.
Motivation & Objective
- Assess whether modern, highly accurate DNNs better model inferotemporal (IT) cortex responses to natural images.
- Investigate why task-optimized DNNs lose alignment with IT as they scale.
- Evaluate whether alternative training routines or biological constraints can improve IT predictivity.
- Propose and test a training routine (neural harmonizer) to align DNN representations with human visual features and IT responses.
Proposed method
- Evaluate 135 diverse DNNs (CNNs, ViTs, self-supervised, robustness-trained) pretrained on ImageNet or other data using Brain-Score style neural prediction.
- Record spatially resolved IT neuron responses from two monkeys to high-resolution natural images.
- Train and test DNNs with the neural harmonizer to align human feature importance maps with DNN representations.
- Use partial least squares regression to map DNN unit activity to IT neuron responses and compute neural predictivity.
- Apply CRAFT-based feature decomposition to interpret which image features drive IT responses in harmonized vs standard DNNs.
- Compare neural predictivity across models and time bins, and assess feature alignment between IT and DNNs.

Experimental results
Research questions
- RQ1Does higher ImageNet accuracy correlate with better IT neural predictivity across modern DNNs?
- RQ2What features do ImageNet-trained DNNs rely on, and how do these differ from IT encoding of natural images?
- RQ3Can aligning DNN representations with human perceptual features (neural harmonizer) improve IT predictivity without sacrificing accuracy?
- RQ4Do biologically aligned training routines mitigate the mismatch between object recognition and neural data?
Key findings
- DNNs pretrained on ImageNet become less accurate at predicting IT neuron responses as ImageNet accuracy increases.
- Training methods like self-supervision or adversarial robustness do not resolve the IT prediction trade-off.
- Harmonized DNNs (hDNNs) significantly improve IT predictivity across PL and ML regions in two monkeys.
- Harmonized models reveal features driving IT activity that align with human judgments (e.g., facial parts) rather than background features.
- CRAFT-based analysis shows harmonized models provide testable, interpretable hypotheses about IT feature selectivity.
- Harmonized models break the pareto-front between ImageNet accuracy and neural prediction accuracy observed in standard DNNs.
![Figure 2 : IT recordings that reveal spatial maps of neuronal responses to complex natural images offer unprecedented insights into their feature selectivity [ 6 ] . (a) Neurons in posterior (PL) and/or medial (ML) lateral IT in two animals were localized using functional magnetic resonance imaging](https://ar5iv.labs.arxiv.org/html/2306.03779/assets/figures/method.png)
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.