[論文レビュー] A Survey of Embodied AI: From Simulators to Research Tasks
embodied AI の包括的百科事典的調査、九つのシミュレーターをベンチマークし、三つの主要な研究タスク:視覚的探索、視覚ナビゲーション、そして embodyd QA を詳述し、シミュレーターとタスクの適合性および将来の方向性に関する指針を提供する。
There has been an emerging paradigm shift from the era of "internet AI" to "embodied AI", where AI algorithms and agents no longer learn from datasets of images, videos or text curated primarily from the internet. Instead, they learn through interactions with their environments from an egocentric perception similar to humans. Consequently, there has been substantial growth in the demand for embodied AI simulators to support various embodied AI research tasks. This growing interest in embodied AI is beneficial to the greater pursuit of Artificial General Intelligence (AGI), but there has not been a contemporary and comprehensive survey of this field. This paper aims to provide an encyclopedic survey for the field of embodied AI, from its simulators to its research. By evaluating nine current embodied AI simulators with our proposed seven features, this paper aims to understand the simulators in their provision for use in embodied AI research and their limitations. Lastly, this paper surveys the three main research tasks in embodied AI -- visual exploration, visual navigation and embodied question answering (QA), covering the state-of-the-art approaches, evaluation metrics and datasets. Finally, with the new insights revealed through surveying the field, the paper will provide suggestions for simulator-for-task selections and recommendations for the future directions of the field.
研究の動機と目的
- 視覚的探索、視覚ナビゲーション、 embodied QA の三つの核心タスクにおける研究の現状を、シミュレーターから研究タスクへと発展させることを調査する。
- 現実性、スケーラビリティ、対話性の観点から九つの embodied AI シミュレーターをベンチマークする。
- 三つの核心タスクに対する最先端のアプローチ、評価指標、およびデータセットを要約する:視覚的探索、視覚ナビゲーション、embodied QA。
- 特定の研究タスクに対するシミュレーター選択のガイダンスを提供し、将来の方向性を提案する。
提案手法
- シミュレーターを評価するために用いられる七つの技術的特徴:Environment、Physics、Object Type、Object Property、Controller、Action、and Multi-Agent.
- 現実性、スケーラビリティ、対話性に基づく二次的評価特徴。
- 七つの特徴にわたるシミュレーター(Table I and II)について、定性的および定量的比較を網羅する。
- 三つの主要な embodiment AI 研究タスクと、それらの最先端の方法論、評価指標、データセット(Table III)の調査。
- シミュレーター、データセット、タスク間の相互関係を分析し、課題を特定する。
実験結果
リサーチクエスチョン
- RQ1現実性、スケーラビリティ、対話性の点で、現在の embodiment AI シミュレーターの能力と制限は何か。
- RQ2視覚的探索、視覚ナビゲーション、embodied QA などのコアタスクを、さまざまなシミュレーターはどのようにサポートしているか。
- RQ3特定の embodiment AI タスクに適したシミュレーターとデータセットを選択する際、研究者を支援するガイドラインは何か。
- RQ4embodiment AI 研究およびシミュレーションフレームワークの主要な課題と将来の方向性は何か。
主な発見
- 九つの embodiment AI シミュレーター(DeepMind Lab, AI2-THOR, CHALET, VirtualHome, VRKitchen, Habitat-Sim, iGibson, SAPIEN, ThreeDWorld)は七つの特徴でベンチマークされている。
- 現実性、スケーラビリティ、対話性を、シミュレーターを比較する二次的評価特徴として提案する。
- AI2-THOR、iGibson、Habitat-Sim は広範な現実性、対話性、スケーラビリティを提供し、多様な embodiment AI タスクに人気がある。
- 三つの主要タスク—視覚的探索、視覚ナビゲーション、embodied QA—は、最先端のアプローチ、評価指標、データセットをカバーしている。
- 本論文はシミュレーターとタスクの選択に関する推奨と、将来の研究方向を提供する。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。