[Paper Review] A Survey of Embodied AI: From Simulators to Research Tasks
A comprehensive encyclopedic survey of embodied AI, benchmarking nine simulators and detailing three main research tasks: visual exploration, visual navigation, and embodied QA, with guidance on simulator-task matching and future directions.
There has been an emerging paradigm shift from the era of "internet AI" to "embodied AI", where AI algorithms and agents no longer learn from datasets of images, videos or text curated primarily from the internet. Instead, they learn through interactions with their environments from an egocentric perception similar to humans. Consequently, there has been substantial growth in the demand for embodied AI simulators to support various embodied AI research tasks. This growing interest in embodied AI is beneficial to the greater pursuit of Artificial General Intelligence (AGI), but there has not been a contemporary and comprehensive survey of this field. This paper aims to provide an encyclopedic survey for the field of embodied AI, from its simulators to its research. By evaluating nine current embodied AI simulators with our proposed seven features, this paper aims to understand the simulators in their provision for use in embodied AI research and their limitations. Lastly, this paper surveys the three main research tasks in embodied AI -- visual exploration, visual navigation and embodied question answering (QA), covering the state-of-the-art approaches, evaluation metrics and datasets. Finally, with the new insights revealed through surveying the field, the paper will provide suggestions for simulator-for-task selections and recommendations for the future directions of the field.
Motivation & Objective
- Survey the development of embodied AI from simulators to research tasks.
- Benchmark nine embodied AI simulators across realism, scalability, and interactivity.
- Summarize state-of-the-art approaches, evaluation metrics, and datasets for three core tasks: visual exploration, visual navigation, and embodied QA.
- Provide guidance on simulator selection for specific research tasks and propose future directions.
Proposed method
- Seven technical features used to evaluate simulators: Environment, Physics, Object Type, Object Property, Controller, Action, and Multi-Agent.
- Secondary evaluation features based on realism, scalability, and interactivity.
- Comprehensive qualitative and quantitative comparison of simulators (Table I and II) across the seven features.
- Survey of three main embodied AI research tasks and their state-of-the-art methods, metrics, and datasets (Table III).
- Analysis of interconnections between simulators, datasets, and tasks to identify challenges.
Experimental results
Research questions
- RQ1What are the capabilities and limitations of current embodied AI simulators for realism, scalability, and interactivity?
- RQ2How do different simulators support core embodied AI tasks such as visual exploration, visual navigation, and embodied QA?
- RQ3What guidelines can help researchers select appropriate simulators and datasets for specific embodied AI tasks?
- RQ4What are the key challenges and future directions in embodied AI research and simulation frameworks.
Key findings
- Nine embodied AI simulators (DeepMind Lab, AI2-THOR, CHALET, VirtualHome, VRKitchen, Habitat-Sim, iGibson, SAPIEN, ThreeDWorld) are benchmarked on seven features.
- Realism, scalability, and interactivity are proposed as three secondary evaluation features to compare simulators.
- AI2-THOR, iGibson, and Habitat-Sim offer broad realism, interactivity, and scalability, making them popular for diverse embodied AI tasks.
- The three main tasks—visual exploration, visual navigation, and embodied QA—cover state-of-the-art approaches, evaluation metrics, and datasets.
- The paper provides recommendations for simulator-task selections and directions for future research.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.