Skip to main content
QUICK REVIEW

[論文レビュー] A Dual-Stream Neural Network Explains the Functional Segregation of Dorsal and Ventral Visual Pathways in Human Brains

Minkyu Choi, Kuan Han|arXiv (Cornell University)|Oct 20, 2023
Visual perception and processing mechanisms被引用数 6
ひとこと要約

本論文は、人間の後頭側(どこ)および腹側(何)視覚経路を模倣する脳にインspiredされた二重ストリームニューラルネットワークを提案する。このモデルは、別々の網膜サンプリングと、空間的注意(後頭側ストリーム)と物体認識(腹側ストリーム)のための異なる学習目的を用いる。注意駆動型の目動かしとストリーム間の再帰的相互作用を活用することで、機能的分離が主に入力の違いではなく、目的の違いに起因するという点で、人間の脳応答との優れた機能的整合性を達成する。

ABSTRACT

The human visual system uses two parallel pathways for spatial processing and object recognition. In contrast, computer vision systems tend to use a single feedforward pathway, rendering them less robust, adaptive, or efficient than human vision. To bridge this gap, we developed a dual-stream vision model inspired by the human eyes and brain. At the input level, the model samples two complementary visual patterns to mimic how the human eyes use magnocellular and parvocellular retinal ganglion cells to separate retinal inputs to the brain. At the backend, the model processes the separate input patterns through two branches of convolutional neural networks (CNN) to mimic how the human brain uses the dorsal and ventral cortical pathways for parallel visual processing. The first branch (WhereCNN) samples a global view to learn spatial attention and control eye movements. The second branch (WhatCNN) samples a local view to represent the object around the fixation. Over time, the two branches interact recurrently to build a scene representation from moving fixations. We compared this model with the human brains processing the same movie and evaluated their functional alignment by linear transformation. The WhereCNN and WhatCNN branches were found to differentially match the dorsal and ventral pathways of the visual cortex, respectively, primarily due to their different learning objectives. These model-based results lead us to speculate that the distinct responses and representations of the ventral and dorsal streams are more influenced by their distinct goals in visual attention and object recognition than by their specific bias or selectivity in retinal inputs. This dual-stream model takes a further step in brain-inspired computer vision, enabling parallel neural networks to actively explore and understand the visual surroundings.

研究の動機と目的

  • 人間の視覚とコンピュータビジョンのギャップを埋めるために、脳の二重ストリーム視覚処理アーキテクチャをモデル化すること。
  • 後頭側および腹側視覚経路の機能的分離が、網膜入力の違いではなく、異なる学習目的に起因するかどうかを調査すること。
  • 能動的で注意駆動型の視覚的探索を模倣することで、コンピュータビジョンシステムのロバスト性と適応性を向上させること。
  • 注意駆動型の目動かしが、人間の視覚処理のニューラルネットワークモデルにおける符号化性能をどの程度向上させるかを評価すること。

提案手法

  • モデルは、後頭側ストリームに適した広範囲/粗いサンプリングと、腹側ストリームに適した狭範囲/細かいサンプリングの2つの補完的網膜サンプリングパターンを用い、マグノセルラーやパラヴォキセルラーカラムを模倣する。
  • WhereCNNは、空間的注意を用いてグローバルな視界を処理し、注視点の位置を予測して誘導する。
  • WhatCNNは、注視点周辺のローカルな視界を処理し、時間経過とともにシーン全体の理解を再構築する物体表現を構築する。
  • 2つのストリームは、逐次的な注視点移動を通じて再帰的に相互作用し、動的なシーン表現を可能にする。
  • 機能的整合性は、後頭側および腹側視覚領域におけるモデルの予測と脳応答を比較するための線形変換を用いて、人間のfMRIデータと評価される。
  • モデルの性能は、単一ストリームネットワーク(例:AlexNet、ResNet)およびランダムな注視点ベースラインと比較され、注意と二重ストリーム設計の影響を隔離する。

実験結果

リサーチクエスチョン

  • RQ1人間の視覚系における機能的分離は、主に網膜入力サンプリングの違いではなく、異なる学習目的に起因するか?
  • RQ2注意駆動型の目動かしは、モデルが人間の脳応答を予測する能力をどの程度向上させるか?
  • RQ3二重ストリームアーキテクチャは、単一ストリームモデルと比較して、後頭側および腹側視覚皮質領域の応答を予測する上でどのように異なるか?
  • RQ4ストリーム間の再帰的相互作用は、シーン表現の向上と脳との機能的整合性の向上に寄与するか?

主な発見

  • WhereCNNとWhatCNNは、それぞれ後頭側および腹側視覚経路と差別的な機能的整合性を示したが、これは主に学習目的の違いに起因し、網膜サンプリングの違いによるものではなかった。
  • モデルは、AlexNetのような単一ストリームネットワークよりも優れた符号化性能を達成し、ResNet18やResNet34のようなより深いモデルと同等の性能を示した。
  • 注意駆動型の注視点は、ほとんどの視覚皮質領域で予測精度を顕著に向上させ、両ストリームの温色のボクセル(Δr > 0)において高い符号化性能を示した。
  • 学習された注視点の使用は、特に後頭側および腹側ストリーム領域において、ランダムな注視点よりも正確な脳応答の予測を可能にした。
  • 二重ストリームモデルは、独立した単一ストリームモデルを上回った。これは、相互作用的で並列的な処理が、人間の脳との機能的整合性を向上させることを示している。
  • 結果から、人間の脳における後頭側-腹側の機能的分離は、感覚的入力バイアスよりも、空間的注意と物体認識という異なる認知的目標に強く依存していると考えられる。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。