Skip to main content
QUICK REVIEW

[論文レビュー] Robustified ANNs Reveal Wormholes Between Human Category Percepts

Guy Gaziv, Michael J. Lee|arXiv (Cornell University)|Aug 14, 2023
Neural dynamics and brain functionNeuroscience被引用数 3
ひとこと要約

本研究では、強化された人工ニューラルネットワーク(ANN)が、人間の物体カテゴリ認識に著しく影響を与える低ノルムの画像摂動を特定できることを示している—これは、このような摂動下でも人間の知覚が安定しているという仮定に疑問を呈するものである。これらの摂動は、画像空間における意味的に異なる知覚状態を接続する「ワームホール」として機能し、従来検出されなかった人間の視覚処理における脆弱性を明らかにしている。このような脆弱性は、最先端のANNを用いることで予測可能である。

ABSTRACT

The visual object category reports of artificial neural networks (ANNs) are notoriously sensitive to tiny, adversarial image perturbations. Because human category reports (aka human percepts) are thought to be insensitive to those same small-norm perturbations -- and locally stable in general -- this argues that ANNs are incomplete scientific models of human visual perception. Consistent with this, we show that when small-norm image perturbations are generated by standard ANN models, human object category percepts are indeed highly stable. However, in this very same "human-presumed-stable" regime, we find that robustified ANNs reliably discover low-norm image perturbations that strongly disrupt human percepts. These previously undetectable human perceptual disruptions are massive in amplitude, approaching the same level of sensitivity seen in robustified ANNs. Further, we show that robustified ANNs support precise perceptual state interventions: they guide the construction of low-norm image perturbations that strongly alter human category percepts toward specific prescribed percepts. These observations suggest that for arbitrary starting points in image space, there exists a set of nearby "wormholes", each leading the subject from their current category perceptual state into a semantically very different state. Moreover, contemporary ANN models of biological visual processing are now accurate enough to consistently guide us to those portals.

研究の動機と目的

  • 人間の物体カテゴリ認識が低ノルムの画像摂動に対して頑健であるという一般的な仮定に反論すること。
  • 標準的なモデルでは検出できない人間の知覚的混乱が、強化されたANNによって明らかにできるかどうかを調査すること。
  • 知覚が安定するとされる低ピクセル予算領域において、人間の知覚を正確に、標的的に変調することが可能かどうかを特定すること。
  • 霊長目腹側ストリームの現代的ANNモデルが、これらの知覚的「ワームホール」の発見を的確に導けるかどうかを評価すること。

提案手法

  • 小規模な摂動に対して耐性を持つよう、ℓ₂-ノルム制約付きの訓練を用いてResNet50ベースのANNを敵対的頑健化すること。
  • 強化されたANNの潜在空間において勾配ベースの最適化を用いて、特定のカテゴリシフトを狙った低ノルムの画像摂動を生成すること。
  • 模擬的な人間の分類行動を再現するための「サーヴィレート」モデルを用い、知覚的混乱の効果を検証すること。
  • 119名の被験者に対して摂動を加えた画像の知覚報告を収集し、カテゴリシフト率とミス率を測定すること。
  • ミス率補正を適用して、ミスがゼロの状態における人間の行動を推定し、知覚的整合性指標の信頼性を向上させること。
  • 通常のANNと強化されたANNを、低ノルム摂動(≤30 ℓ₂-ノルム)下での人間の知覚的シフト予測において体系的に比較すること。
Figure 1: Robustified models discover low-norm image perturbations that strongly modulate human category percepts. The prevailing assumption: Human object category percepts have complicated topology in pixel space, but are robust (i.e., stable) inside a low pixel budget envelope around most natural
Figure 1: Robustified models discover low-norm image perturbations that strongly modulate human category percepts. The prevailing assumption: Human object category percepts have complicated topology in pixel space, but are robust (i.e., stable) inside a low pixel budget envelope around most natural

実験結果

リサーチクエスチョン

  • RQ1強化されたANNは、低ノルム摂動領域における人工知覚と人間知覚の行動的整合性のギャップを埋めることができるか?
  • RQ2強化されたANNが生成した低ノルム画像摂動は、人間の物体カテゴリ認識に強く、かつ信頼性のある変化を引き起こすことができるか?
  • RQ3画像空間全体にわたり、ある知覚状態から意味的に離れた別の状態へとつながる『ワームホール』—近接する画像摂動—が存在するか?
  • RQ4強化されたANNは、人間の知覚を任意の所定のカテゴリに向けて正確に標的にして変調できるか?

主な発見

  • 強化されたANNは、人間被験者に最大約90%の割合でカテゴリの混乱を引き起こす低ノルムの画像摂動(≤30 ℓ₂-ノルム)を発見し、知覚の頑健性という仮定に疑問を呈した。
  • 強化されたANNによって誘導された摂動では、人間の知覚反応が特定のカテゴリに向けて強く標的的に変調され、マルチターゲット変調タスクにおいて約60%の混乱率を示した。
  • 通常のANNと人間の間には大きな行動的整合性のギャップが存在したが、強化されたANNは特に低ピクセル予算領域においてこのギャップを顕著に縮小した。
  • 強化されたANNは、人間の知覚を一から別のカテゴリに信頼性高くシフトできるような、正確な低ノルム干渉を可能にした。これは、知覚空間に『ワームホール』が存在することを示唆している。
  • 結果から、霊長目の腹側ストリームを模倣する現代のANNモデルは、これらの知覚的ポータルの発見を一貫して導けるほど十分に正確であることが示唆された。
  • 強い整合性が得られたにもかかわらず、依然として残存する行動的ギャップが存在し、強化されたANNですら人間の視覚処理のすべての側面を完全に捉えきれていない可能性を示唆している。
Figure 2: Low-norm image perturbations discovered by robustified models strongly disrupt human category judgements. (a) The Guide Models used for Disruption Modulation (DM) image generation. (b) Disruption rates of humans and models. All panels share the same set of start images, and the four sets o
Figure 2: Low-norm image perturbations discovered by robustified models strongly disrupt human category judgements. (a) The Guide Models used for Disruption Modulation (DM) image generation. (b) Disruption rates of humans and models. All panels share the same set of start images, and the four sets o

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。