[论文解读] A Survey of Embodied AI: From Simulators to Research Tasks
一份全面的、百科式的具身 AI 调查,基准测试九个仿真器并详细阐述三个主要研究任务:视觉探索、视觉导航,以及 embodied QA,附带关于仿真器与任务匹配的指南及未来方向。
There has been an emerging paradigm shift from the era of "internet AI" to "embodied AI", where AI algorithms and agents no longer learn from datasets of images, videos or text curated primarily from the internet. Instead, they learn through interactions with their environments from an egocentric perception similar to humans. Consequently, there has been substantial growth in the demand for embodied AI simulators to support various embodied AI research tasks. This growing interest in embodied AI is beneficial to the greater pursuit of Artificial General Intelligence (AGI), but there has not been a contemporary and comprehensive survey of this field. This paper aims to provide an encyclopedic survey for the field of embodied AI, from its simulators to its research. By evaluating nine current embodied AI simulators with our proposed seven features, this paper aims to understand the simulators in their provision for use in embodied AI research and their limitations. Lastly, this paper surveys the three main research tasks in embodied AI -- visual exploration, visual navigation and embodied question answering (QA), covering the state-of-the-art approaches, evaluation metrics and datasets. Finally, with the new insights revealed through surveying the field, the paper will provide suggestions for simulator-for-task selections and recommendations for the future directions of the field.
研究动机与目标
- 从仿真器到研究任务,调查具身 AI 的发展。
- 在现实性、可扩展性和交互性方面对九个具身 AI 仿真器进行基准测试。
- 总结三个核心任务的最新方法、评估指标和数据集:视觉探索、视觉导航,以及具身 QA。
- 就特定研究任务提供仿真器选择的指导,并提出未来方向。
提出的方法
- 七个用于评估仿真器的技术特征:环境、物理、对象类型、对象属性、控制器、动作,以及多代理。
- 基于现实性、可扩展性和交互性的次要评估特征。
- 在这七个特征上对仿真器进行全面的定性与定量比较(Table I 与 Table II)。
- 对三大具身 AI 研究任务及其最先进的方法、指标和数据集进行综述(Table III)。
- 分析仿真器、数据集和任务之间的联系,以发现挑战。
实验结果
研究问题
- RQ1当前具身 AI 仿真器在现实性、可扩展性和交互性方面的能力与局限性有哪些?
- RQ2不同仿真器如何支持核心具身 AI 任务,如视觉探索、视觉导航和具身 QA?
- RQ3有哪些指南可以帮助研究者为特定具身 AI 任务选择合适的仿真器与数据集?
- RQ4具身 AI 研究与仿真框架的关键挑战与未来方向有哪些。
主要发现
- 九个具身 AI 仿真器(DeepMind Lab, AI2-THOR, CHALET, VirtualHome, VRKitchen, Habitat-Sim, iGibson, SAPIEN, ThreeDWorld)在七个特征上进行了基准测试。
- 将现实性、可扩展性和交互性提议为比较仿真器的三个次要评估特征。
- AI2-THOR、iGibson 和 Habitat-Sim 提供广泛的现实性、交互性和可扩展性,使它们在多样化具身 AI 任务中受到欢迎。
- 三大任务——视觉探索、视觉导航,以及具身 QA——涵盖了前沿方法、评估指标和数据集。
- 论文为仿真器-任务选择提供建议,并提出未来研究方向。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。