Skip to main content
QUICK REVIEW

[論文レビュー] Environmental drivers of systematicity and generalization in a situated agent

Felix Hill, Andrew K. Lampinen|arXiv (Cornell University)|Oct 1, 2019
Language and cultural evolution参考文献 39被引用数 53
ひとこと要約

この論文は、 embodiment(具現化)された多モーダルエージェントが、構成的な言語 grounded な振る舞いをゼロショットで一般化できること、そして一般化が訓練データの多様性、視点制約、現実的な環境における知覚の豊かさによって形成されることを示しています。

ABSTRACT

The question of whether deep neural networks are good at generalising beyond their immediate training experience is of critical importance for learning-based approaches to AI. Here, we consider tests of out-of-sample generalisation that require an agent to respond to never-seen-before instructions by manipulating and positioning objects in a 3D Unity simulated room. We first describe a comparatively generic agent architecture that exhibits strong performance on these tests. We then identify three aspects of the training regime and environment that make a significant difference to its performance: (a) the number of object/word experiences in the training set; (b) the visual invariances afforded by the agent's perspective, or frame of reference; and (c) the variety of visual input inherent in the perceptual aspect of the agent's perception. Our findings indicate that the degree of generalisation that networks exhibit can depend critically on particulars of the environment in which a given task is instantiated. They further suggest that the propensity for neural networks to generalise in systematic ways may increase if, like human children, those networks have access to many frames of richly varying, multi-modal observations as they learn.

研究の動機と目的

  • 標準的なニューラルアーキテクチャが、多モーダルで situated な設定において系統的一般化を達成できるかを調査する。
  • 環境要因が grounded language タスクにおける動詞・名詞の組み合わせ理解の出現に影響を与えるかを判断する。
  • unseen の物体と動作へのゼロショット一般化を高める訓練 regime の要因を特定する。

提案手法

  • 視覚(ピクセル)と言語入力を持つ多モーダルエージェントが、観測を3層CNNとLSTMベースの言語モジュールで処理する。
  • LSTMベースの方策と価値ネットワークが、分散型アクター(IMPALAスタイル)で訓練されたアクタークリティック枠組みの中で動作する。
  • 実験は、物体を持ち上げる・置くゼロショット一般化、述語を引数に束縛する、さまざまな環境条件下での一般化を評価する。
  • 3DのUnity環境と2Dのグリッド世界、自己中心的視点と共立視点、言語 supervision の有無の比較分析。
  • 制御条件として、色・形状一般化タスクで視覚と言語分類器を用いた場合と、 situated エージェントを比較する。

実験結果

リサーチクエスチョン

  • RQ1標準的なニューラルアーキテクチャは、3Dインタラクティブ環境において動詞と名詞を novel object に grounding し、一般化できるか。
  • RQ2環境と知覚の要因は、 situated エージェントの系統的一般化を促進または妨げるか。
  • RQ3訓練の多様性(語・物体)、参照枠の制約、より豊かな時間的知覚を増やすことで、ゼロショット一般化が向上するか。
  • RQ4言語 supervision は grounded タスクの系統的一般化にどの程度寄与するか。

主な発見

  • エージェントは、ゼロショットのリフティング・置作業において novel object および新規語彙と object の組み合わせへ一般化する。
  • 訓練中に経験する語・物体の多様性を増やすと一般化が改善する(否定実験)。
  • 自己中心的/制約付きの視覚的視点は、共立視点と比べて一般化を高める。
  • 時間的に豊かな知覚入力( varied views を移動すること)は、単一フレーム知覚を超える一般化を促す。
  • 3D環境では、同等の2Dグリッド世界設定より一般化が強く、言語は訓練性能に控えめに寄与するが、一般化には必須ではない。
  • 静止画像上の vision-language classifier は、 test generalization で situated エージェントを下回り、 embodiment による知覚の利点を強調する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。