Skip to main content
QUICK REVIEW

[論文レビュー] Atari-HEAD: Atari Human Eye-Tracking and Demonstration Dataset

Ruohan Zhang, Calen Walshe|arXiv (Cornell University)|Mar 15, 2019
Gaze Tracking and Assistive Technology参考文献 42被引用数 14
ひとこと要約

本論文は、20種類のAtari 2600ゲームから収集した人間の眼球追跡データと操作データを含む大規模データセットであるAtari-HEADを紹介する。本研究では、人間の注目(注視)を政策ネットワークに統合することで、強化学習の模倣学習を向上させる注視予測ネットワークを提案し、同じベンチマークゲームにおいて標準的な行動コーディング(behavior cloning)と比較して平均スコアを115.26%向上させた。

ABSTRACT

Large-scale public datasets have been shown to benefit research in multiple areas of modern artificial intelligence. For decision-making research that requires human data, high-quality datasets serve as important benchmarks to facilitate the development of new methods by providing a common reproducible standard. Many human decision-making tasks require visual attention to obtain high levels of performance. Therefore, measuring eye movements can provide a rich source of information about the strategies that humans use to solve decision-making tasks. Here, we provide a large-scale, high-quality dataset of human actions with simultaneously recorded eye movements while humans play Atari video games. The dataset consists of 117 hours of gameplay data from a diverse set of 20 games, with 8 million action demonstrations and 328 million gaze samples. We introduce a novel form of gameplay, in which the human plays in a semi-frame-by-frame manner. This leads to near-optimal game decisions and game scores that are comparable or better than known human records. We demonstrate the usefulness of the dataset through two simple applications: predicting human gaze and imitating human demonstrated actions. The quality of the data leads to promising results in both tasks. Moreover, using a learned human gaze model to inform imitation learning leads to an 115\% increase in game performance. We interpret these results as highlighting the importance of incorporating human visual attention in models of decision making and demonstrating the value of the current dataset to the research community. We hope that the scale and quality of this dataset can provide more opportunities to researchers in the areas of visual attention, imitation learning, and reinforcement learning.

研究の動機と目的

  • 視覚的ダイナミクスやタスク構造が異なる多様なAtari 2600ゲームにおいて、高品質な人間の注視と行動データを収集すること。
  • 複雑で視覚的に豊かな環境における人間の視覚的注目(注視)と意思決定の相関関係を解明すること。
  • 学習された人間の注視パターンを政策ネットワークに組み込むことで、模倣学習のパフォーマンスを向上させること。
  • 報酬が密集している環境と疎らな環境の両方で、注視誘導型模倣学習が標準的な行動コーディングを上回るかを評価すること。

提案手法

  • ゲームプレイ中に眼動-trackingハードウェアを用いて、20種類のAtari 2600ゲームにおける人間プレーヤーの注視と行動データを収集した。
  • 各ゲームごとに、スタックされたグレースケールフレームから注視の注目度マップを予測する専用の3層畳み込みおよび3層デコンボリューションネットワークを訓練した。
  • 予測された注視の注目度マップを、二本のブランチを持つ政策ネットワークのモダリティとして統合した:一方のブランチは生画像を処理し、もう一方は予測された注視でマスクされた画像を処理した。
  • 両ブランチからの特徴量を平均化して行動確率を生成し、注視に配慮した模倣学習(AGIL)を実現した。
  • ネットワーク出力確率に基づいて行動をサンプリングするために、温度パrameter η = 1 のボルツマン探索方策を用いた。
  • 1ゲームあたり500エピソードの平均ゲームスコアを用いてパフォーマンスを評価し、1エピソードのフレーム制限を108Kに設定した。

実験結果

リサーチクエスチョン

  • RQ1深層ネットワークは、多様なAtari 2600ゲームにおいて人間の注視パターンをどれほど正確に予測できるか?
  • RQ2予測された人間の注視を組み込むことで、標準的な行動コーディングと比較して模倣学習のパフォーマンスはどの程度向上するか?
  • RQ3注視誘導型模倣学習は、視覚的複雑さや報酬の疎らさが異なるゲーム間で一般化可能か?
  • RQ4異なるデータセットやアーキテクチャを用いた先行研究と比較して、注視拡張型模倣学習のパフォーマンスはいかがなっているか?

主な発見

  • 注視予測ネットワークは全20ゲームで高い精度を達成し、AUC > 0.94 かつ NSS > 4.0 を達成。ランダムベースライン(AUC = 0.500)およびボトムアップ注視モデルを著しく上回った。
  • AGILエージェントは、標準的な行動コーディング(AtariHead-IL)と比較して、平均して115.26%のスコア向上を達成し、スコア上昇幅は+0.003%(freeway)から+112.33%(alien)の範囲であった。
  • モンテズマのレインジやMs. パックマンのようなゲームでは、それぞれ88.8%および67.8%の成功率を達成し、標準的なIL(86.6%および55.5%)を著しく上回った。
  • 注視拡張型ポリシーは20ゲーム中18ゲームでパフォーマンス向上を達成し、特に報酬が疎らなゲーム(Seaquest:+309.05%、Frostbite:+52.05%)で最大の向上を示した。
  • 注視ネットワークはゲーム全体で平均AUC 0.973を達成し、個別スコアはMs. パックマンで0.945、Enduroで0.988の範囲であった。
  • ロードランナーやスペースインベーダーズのようなゲームでは、AGILエージェントはそれぞれ42,539.4点および248.2点を達成し、Kurin-ILおよびHester-ILのベースラインを上回った。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。