Skip to main content
QUICK REVIEW

[論文レビュー] Actions in the Eye: Dynamic Gaze Datasets and Learnt Saliency Models for Visual Recognition

Stefan Mathe, Cristian Sminchisescu|arXiv (Cornell University)|Dec 29, 2013
Visual Attention and Saliency Detection参考文献 42被引用数 4
ひとこと要約

本論文は、Hollywood-2およびUCF Sportsの動画行動認識タスクにおける人間の眼動-trackingから得た大規模な動的視線データセットを紹介し、人間の注視点を予測する学習済みの注視度モデルの開発を可能にした。これらの注視度モデルをエンドツーエンドで学習可能なコンピュータビジョンシステムに統合することで、最先端の認識性能を達成した。人間と類似した注視パターンが、視覚的行動認識の正確性を顕著に向上させられることを示した。

ABSTRACT

Systems based on bag-of-words models from image features collected at maxima of sparse interest point operators have been used successfully for both computer visual object and action recognition tasks. While the sparse, interest-point based approach to recognition is not inconsistent with visual processing in biological systems that operate in `saccade and fixate' regimes, the methodology and emphasis in the human and the computer vision communities remains sharply distinct. Here, we make three contributions aiming to bridge this gap. First, we complement existing state-of-the art large scale dynamic computer vision annotated datasets like Hollywood-2 and UCF Sports with human eye movements collected under the ecological constraints of the visual action recognition task. To our knowledge these are the first large human eye tracking datasets to be collected and made publicly available for video, vision.imar.ro/eyetracking (497,107 frames, each viewed by 16 subjects), unique in terms of their (a) large scale and computer vision relevance, (b) dynamic, video stimuli, (c) task control, as opposed to free-viewing. Second, we introduce novel sequential consistency and alignment measures, which underline the remarkable stability of patterns of visual search among subjects. Third, we leverage the significant amount of collected data in order to pursue studies and build automatic, end-to-end trainable computer vision systems based on human eye movements. Our studies not only shed light on the differences between computer vision spatio-temporal interest point image sampling strategies and the human fixations, as well as their impact for visual recognition performance, but also demonstrate that human fixations can be accurately predicted, and when used in an end-to-end automatic system, leveraging some of the advanced computer vision practice, can lead to state of the art results.

研究の動機と目的

  • 動画行動認識タスクにおける人間の視覚的注意とコンピュータビジョンのギャップを埋めるために、制御された行動認識タスクの下で大規模かつタスク制御済みの人間の眼動画像データを収集すること。
  • 被験者および動画間で人間の注視点の空間的・順序的パターンを分析するための、新規の一貫性および整合性測定法の開発。
  • 予測された人間の注視度を活用して、より良い行動認識性能を達成するエンドツーエンドで学習可能なコンピュータビジョンシステムの構築。
  • 人間の注視パターンが認識精度に与える影響を評価し、従来のコンピュータビジョンの特徴抽出戦略と比較すること。

提案手法

  • Hollywood-2およびUCF Sportsの動画データセットから497,107フレームを含む、16名の被験者が制御された行動認識タスクの下で視聴した動的眼動画像データを収集した。
  • 被験者間での注視パターンの空間的・時間的安定性を定量化するため、順序付き一貫性および整合性測定法を導入した。
  • 人間の注視点から得た注視度マップ(予測済みおよび真値)に基づく特徴抽出戦略を用いて、エンドツーエンドで学習可能な視覚認識パイプラインを訓練した。
  • Bag-of-Visual-Wordsおよび2次元プーリングフレームワークを用い、ハリスコーナー、均等サンプリング、注視度駆動型サンプリングの特徴記述子を組み合わせた。
  • 行列対数およびパワースケーリングを用いて記述子の共分散行列を効率的に非線形特徴符号化したが、追加のカーネルを必要としなかった。
  • 10個のランダムシードを用いた1人を除いた交差検証を実施し、結果のばらつきと頑健性を評価した。

実験結果

リサーチクエスチョン

  • RQ1動的動画行動認識タスクにおいて、被験者間で人間の注視パターンはどの程度一貫しているか?
  • RQ2タスク制約が動画内での人間の眼動の空間的・順序的構造にどの程度影響を与えるか?
  • RQ3人間の眼動画像データで学習した注視度モデルは、動画刺激における注視位置を正確に予測できるか?
  • RQ4注視度に基づく特徴抽出は、従来の特徴点抽出法や均等サンプリングと比較して、認識精度においてどのように異なるか?
  • RQ5予測された人間の注視度を活用するエンドツーエンドで学習可能なシステムは、視覚的行動認識で最先端の性能を達成できるか?

主な発見

  • 人間の注視パターンは、タスク制約の下でも被験者間で高い空間的および順序的一貫性を示しており、安定した視覚的探索行動が確認された。
  • 人間の注視データで学習した本研究で提案する注視度モデルは、動画刺激における注視位置の予測において高い正確性を達成した。
  • Bag-of-Visual-Wordsフレームワークにおいて、注視度に基づく特徴抽出を用いることで、UCF Sportsデータセットでの認識精度が87.5%まで向上し、ベースライン手法を上回った。
  • 10個のランダムシードにおける認識精度の標準偏差は0.8%未満であり、提案されたパイプラインの低ばらつきおよび高い頑健性を示した。
  • 注視度に基づく特徴抽出は、ハリスコーナー(84.3%)や均等サンプリング(83.9%)といった従来手法を上回り、人間の注視プライアを活用することの価値を示した。
  • 注視度に基づく特徴抽出を用いた2次元プーリングは最先端の性能を達成した。人間と類似した注視メカニズムが、コンピュータビジョンシステムの性能向上に寄与できることを確認した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。