Skip to main content
QUICK REVIEW

[論文レビュー] Understanding More about Human and Machine Attention in Deep Neural Networks

Qiuxia Lai, Salman Khan|arXiv (Cornell University)|Jun 20, 2019
Visual Attention and Saliency Detection参考文献 91被引用数 15
ひとこと要約

本論文は、顕著オブジェクトセグメンテーション、動画アクション認識、細分化分類の3つのコンピュータビジョンタスクにおいて、人間の視線注意と人工的注意の整合性を調査する。実際の人間の注視データをベンチマークとして用い、人間の注意に近い注視メカニズムを備えたモデルが、特に低レベルで注意駆動のタスクにおいて、より優れた性能と頑健性を示すことを示している。これは、解釈可能性と性能の向上のため、ネットワーク設計における明示的な整合性の確保を提唱する。

ABSTRACT

Human visual system can selectively attend to parts of a scene for quick perception, a biological mechanism known as Human attention. Inspired by this, recent deep learning models encode attention mechanisms to focus on the most task-relevant parts of the input signal for further processing, which is called Machine/Neural/Artificial attention. Understanding the relation between human and machine attention is important for interpreting and designing neural networks. Many works claim that the attention mechanism offers an extra dimension of interpretability by explaining where the neural networks look. However, recent studies demonstrate that artificial attention maps do not always coincide with common intuition. In view of these conflicting evidence, here we make a systematic study on using artificial attention and human attention in neural network design. With three example computer vision tasks, diverse representative backbones, and famous architectures, corresponding real human gaze data, and systematically conducted large-scale quantitative studies, we quantify the consistency between artificial attention and human visual attention and offer novel insights into existing artificial attention mechanisms by giving preliminary answers to several key questions related to human and artificial attention mechanisms. Overall results demonstrate that human attention can benchmark the meaningful `ground-truth' in attention-driven tasks, where the more the artificial attention is close to human attention, the better the performance; for higher-level vision tasks, it is case-by-case. It would be advisable for attention-driven tasks to explicitly force a better alignment between artificial and human attention to boost the performance; such alignment would also improve the network explainability for higher-level computer vision tasks.

研究の動機と目的

  • 深層ニューラルネットワークにおける人工的注意が、タスク固有の条件下で人間の視覚的注意と整合しているかどうかを調査すること。
  • 人間の注意が、注意駆動型コンピュータビジョンタスクにおける意味のある注意の信頼できる「真値」ベンチマークとして機能するかどうかを評価すること。
  • ネットワークアーキテクチャ、深さ、および注意メカニズム設計が、人工的注意と人間の注意の整合性に与える影響を分析すること。
  • 注意の整合性が、敵対的攻撃に対するモデルの頑健性および予測精度に与える影響を評価すること。
  • 深層学習におけるより解釈可能で効果的な注意メカニズムを設計するための実用的知見を提供すること。

提案手法

  • 研究は、制御された、目的志向のタスク条件下で収集された実際の人間の注視データを、人間の注意の代理として用いる。
  • 3つのビジョンタスクをカバーする複数の深層学習バックボーン(AlexNet、VGGNet、ResNet)およびアーキテクチャ(Two-stream、FCN)を評価する。
  • 訓練済みモデルから人工的注意マップを抽出し、相関や被り具合などの整合性指標を用いて人間の注視マップと定量的に比較する。
  • 活性化関数(シグモイド、ソフトマックス)のアブレーションスタディ、統合戦略(早期統合対後期統合)、人間の注視からの監督を含む実験を実施する。
  • FGSM敵対的攻撃の下での頑健性をテストし、注意がモデル安定性に果たす役割を評価する。
  • 一般化性を確保するため、多様なデータセットおよびネットワーク設定をカバーする大規模かつ体系的な評価を実施する。

実験結果

リサーチクエスチョン

  • RQ1タスク固有の設定下で、深層ニューラルネットワークにおける人工的注意マップは、どの程度人間の視覚的注意と整合しているか?
  • RQ2人間の注意は、注意駆動型タスクにおける人工的注意の意味の有無を評価する信頼できるベンチマークとして機能できるか?
  • RQ3ネットワークアーキテクチャ、深さ、および注意メカニズム設計は、人工的注意と人間の注意の整合性にどのように影響するか?
  • RQ4人工的注意を人間の注意に合わせることで、モデルの性能と敵対的攻撃に対する頑健性が向上するか?
  • RQ5異なるビジョンタスクにおいて、注意の質と予測精度の間に一貫した関係があるか?

主な発見

  • 人間の注視パターンに近い整合性を示す人工的注意メカニズムは、特に顕著オブジェクトセグメンテーションや細分化分類といった低レベルタスクにおいて、顕著に優れた性能を達成する。
  • 人間の注意は、注意駆動型タスクにおける意味のある注意の有効で妥当なベンチマークであり、高い整合性がモデルの正確性向上と相関している。
  • ResNetベースのモデルは、VGGNet や AlexNet よりも一貫して人間の注視に近い注意マップを生成し、人間の注視の監視によりより大きな恩恵を受ける。
  • 活性化関数の選択(例:動画アクション認識ではシグモイド、細分化分類ではソフトマックス)は、注意マップの品質およびモデル性能に顕著な影響を与える。
  • 正しい視覚的詳細への注目は、頑健性を向上させる。人間の注意に整合したモデルは、FGSM敵対的攻撃に対して顕著に高い耐性を示す。
  • 高レベルタスク(動画アクション認識)では、注意の整合性が有益であるが、予測可能性は低く、タスクの複雑さが人間の注意への直接的な整合性を制限していると考えられる。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。