Skip to main content
QUICK REVIEW

[論文レビュー] GAIA-1: A Generative World Model for Autonomous Driving

Anthony Hu, Lloyd Russell|arXiv (Cornell University)|Sep 29, 2023
Human Motion and Animation被引用数 27
ひとこと要約

GAIA-1 は世界モデル・トランスフォーマーと動画拡散デコーダを組み合わせ、マルチモーダルなプロンプトから現実的な運転シナリオを生成、将来予測、シーン理解、ego車の挙動の細かな制御を実現する。

ABSTRACT

Autonomous driving promises transformative improvements to transportation, but building systems capable of safely navigating the unstructured complexity of real-world scenarios remains challenging. A critical problem lies in effectively predicting the various potential outcomes that may emerge in response to the vehicle's actions as the world evolves. To address this challenge, we introduce GAIA-1 ('Generative AI for Autonomy'), a generative world model that leverages video, text, and action inputs to generate realistic driving scenarios while offering fine-grained control over ego-vehicle behavior and scene features. Our approach casts world modeling as an unsupervised sequence modeling problem by mapping the inputs to discrete tokens, and predicting the next token in the sequence. Emerging properties from our model include learning high-level structures and scene dynamics, contextual awareness, generalization, and understanding of geometry. The power of GAIA-1's learned representation that captures expectations of future events, combined with its ability to generate realistic samples, provides new possibilities for innovation in the field of autonomy, enabling enhanced and accelerated training of autonomous driving technology.

研究の動機と目的

  • 多様な条件下で未来の運転イベントを予測するための、スケーラブルで自己教師なしの世界モデルを開発する。
  • 実世界データから道路シーンとダイナミクスの意味のある高レベル表現を学習する。
  • アクションと言語プロンプトを通じて、ego車の挙動とシーン要素の制御可能な生成を可能にする。
  • 長期的なシーン生成、一般化、3D幾何理解といった出現特性を示す。

提案手法

  • システムを世界モデルと動画拡散デコーダに分割し、シーン推論と高品質な映像レンダリングを分離する。
  • 学習済み画像トークナイザーを用いて各動画フレームを離散的な画像トークンで表現し、未来を次トークン予測として連続的にモデル化する。
  • 過去の画像・テキスト・アクショントークンを条件とした次の画像トークンを予測する自己回帰型トランスフォーマーを世界モデルとして用いる。
  • 世界モデルトークンを条件として高解像度動画をレンダリングし、時間的アップサンプリングを行えるマルチタスク動画拡散デコーダを訓練する。
  • 地理や天候を均等にサンプリングした大規模な実世界の英国都市部運転データセットで訓練し、頑健な表現を学習する。
  • マルチモーダルプロンプト(動画、テキスト、アクション)を取り入れ、推論時に分類子なしガイダンスを用いて生成される未来をテキストプロンプトと一致させる。

実験結果

リサーチクエスチョン

  • RQ1GAIA-1 はマルチモーダルプロンプトから妥当な未来の運転シナリオを確実に予測できるか?
  • RQ2学習されたトークンと出現表現は、自動運転に関連する高レベルのシーン構造・幾何・ダイナミクスを捉えているか?
  • RQ3単一の文脈から複数の妥当な未来を生成できるか?
  • RQ4アクションとテキストプロンプトを介して ego車のダイナミクスとシーン要素をどの程度制御できるか?
  • RQ5スケーリング(データと計算資源)が世界モデルの性能とサンプル品質に与える影響はどの程度か?

主な発見

  • GAIA-1 は高レベルの構造とシーンダイナミクスを学習し、一貫性があり妥当な運転シーンを生成できる。
  • モデルは一般化し創造性を示し、学習例を超えた新規の未来を生み出す。
  • 文脈認識および3D幾何の理解を示し、道路によるピッチ/ロール効果を含む。
  • GAIA-1 は想像から長く安定した運転動画を生成し、同じ文脈から複数の妥当な未来を作り出せる。
  • テキストプロンプトとアクションを介した ego車の挙動とシーン属性の細かな制御が可能である。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。