[論文レビュー] Performance-optimized deep neural networks are evolving into worse models of inferotemporal visual cortex
この論文は、DNNがImageNetで向上するにつれて、IT神経応答を予測する能力が低下することを示している; 神経ハーモナイザーで訓練すると表現を人間と整列させ、神経予測性を回復する。
One of the most impactful findings in computational neuroscience over the past decade is that the object recognition accuracy of deep neural networks (DNNs) correlates with their ability to predict neural responses to natural images in the inferotemporal (IT) cortex. This discovery supported the long-held theory that object recognition is a core objective of the visual cortex, and suggested that more accurate DNNs would serve as better models of IT neuron responses to images. Since then, deep learning has undergone a revolution of scale: billion parameter-scale DNNs trained on billions of images are rivaling or outperforming humans at visual tasks including object recognition. Have today's DNNs become more accurate at predicting IT neuron responses to images as they have grown more accurate at object recognition? Surprisingly, across three independent experiments, we find this is not the case. DNNs have become progressively worse models of IT as their accuracy has increased on ImageNet. To understand why DNNs experience this trade-off and evaluate if they are still an appropriate paradigm for modeling the visual system, we turn to recordings of IT that capture spatially resolved maps of neuronal activity elicited by natural images. These neuronal activity maps reveal that DNNs trained on ImageNet learn to rely on different visual features than those encoded by IT and that this problem worsens as their accuracy increases. We successfully resolved this issue with the neural harmonizer, a plug-and-play training routine for DNNs that aligns their learned representations with humans. Our results suggest that harmonized DNNs break the trade-off between ImageNet accuracy and neural prediction accuracy that assails current DNNs and offer a path to more accurate models of biological vision.
研究の動機と目的
- 現代の高精度DNNが自然画像に対する内側頭葉(IT)皮質の反応をより良くモデルできるかを評価する。
- DNNがタスク最適化されるにつれて、ITとの整合性が失われる理由を調査する。
- 代替の訓練方法や生物学的制約がIT予測性を改善できるかを評価する。
- DNN表現と人間の視覚特徴およびIT反応を整合させる訓練ルーチン(神経ハーモナイザー)を提案・検証する。
提案手法
- Brain-Score風スタイルの神経予測を用いて、ImageNetまたは他のデータで事前訓練された135の多様なDNN(CNN、ViT、自己教師付き、頑健性訓練)を評価する。
- 高解像度の自然画像に対する2匹のサルの時空的に分解されたITニューロン応答を記録する。
- 神経ハーモナイザーを用いてDNNを訓練・評価し、人間の特徴重要度マップとDNN表現を整列させる。
- 部分最小二乗回帰を用いてDNNのユニット活動をITニューロン応答に対応付け、神経予測性を算出する。
- CRAFTベースの特徴分解を適用して、ハーモナイズドDNNと標準DNNでIT応答を導く特徴を解釈する。
- モデル間および時間ビンごとの神経予測性を比較し、ITとDNN間の特徴整合を評価する。

実験結果
リサーチクエスチョン
- RQ1高いImageNet精度は、現代のDNN全体でIT神経予測性を向上させるか?
- RQ2ImageNetで訓練されたDNNはどのような特徴に依存しており、これは自然画像のIT符号化とどう異なるか?
- RQ3人間の知覚特徴とDNN表現を整列させる(神経ハーモナイザー)ことで、精度を犠牲にせずIT予測性を向上させられるか?
- RQ4生物学的に整列した訓練ルーチンは、対象物認識と神経データの不一致を緩和するか?
主な発見
- ImageNetで事前訓練したDNNは、ImageNetの精度が高まるにつれてITニューロン応答を予測する精度が低下する。
- 自己教師付きや敵対的頑健性のような訓練方法は、IT予測のトレードオフを解決しない。
- ハーモナイズドDNN(hDNN)は、二匹のサルにおけるPLおよびML領域のIT予測性を大幅に改善する。
- ハーモナイズドモデルは、人間の判断と一致する特徴(例:顔部位)をIT活動に駆動する特徴を明らかにする。
- CRAFTベースの解析は、ハーモナイズされたモデルがIT特徴選択性について実験可能で解釈可能な仮説を提供する。
- ハーモナイズドモデルは、標準DNNで観察されるImageNet精度と神経予測精度のパレートフロントを破る。
![Figure 2 : IT recordings that reveal spatial maps of neuronal responses to complex natural images offer unprecedented insights into their feature selectivity [ 6 ] . (a) Neurons in posterior (PL) and/or medial (ML) lateral IT in two animals were localized using functional magnetic resonance imaging](https://ar5iv.labs.arxiv.org/html/2306.03779/assets/figures/method.png)
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。