[논문 리뷰] A Survey of Embodied AI: From Simulators to Research Tasks
embodied AI에 대한 포괄적 백과사전식 조사로, 아홉 개의 시뮬레이터를 벤치마킹하고 세 가지 주요 연구 과제인 시각 탐색, 시각 내비게이션, 그리고 구현형 QA를 상세히 다루며 시뮬레이터-과제 매칭 및 향후 방향에 대한 가이드를 제공합니다.
There has been an emerging paradigm shift from the era of "internet AI" to "embodied AI", where AI algorithms and agents no longer learn from datasets of images, videos or text curated primarily from the internet. Instead, they learn through interactions with their environments from an egocentric perception similar to humans. Consequently, there has been substantial growth in the demand for embodied AI simulators to support various embodied AI research tasks. This growing interest in embodied AI is beneficial to the greater pursuit of Artificial General Intelligence (AGI), but there has not been a contemporary and comprehensive survey of this field. This paper aims to provide an encyclopedic survey for the field of embodied AI, from its simulators to its research. By evaluating nine current embodied AI simulators with our proposed seven features, this paper aims to understand the simulators in their provision for use in embodied AI research and their limitations. Lastly, this paper surveys the three main research tasks in embodied AI -- visual exploration, visual navigation and embodied question answering (QA), covering the state-of-the-art approaches, evaluation metrics and datasets. Finally, with the new insights revealed through surveying the field, the paper will provide suggestions for simulator-for-task selections and recommendations for the future directions of the field.
연구 동기 및 목표
- 시뮬레이터에서 연구 과제로 이어지는 구현형 AI의 발전을 조망한다.
- 현실성, 확장성, 상호작용성을 기준으로 아홉 가지 구현형 AI 시뮬레이터를 벤치마크한다.
- 시각 탐색, 시각 내비게이션, 구현형 QA의 세 가지 핵심 작업에 대한 최첨단 방법, 평가 지표 및 데이터 세트를 요약한다.
- 특정 연구 과제에 맞춘 시뮬레이터 선택에 대한 가이드를 제공하고 향후 방향을 제안한다.
제안 방법
- 시뮬레이터를 평가하는 데 사용된 일곱 가지 기술적 특징: Environment, Physics, Object Type, Object Property, Controller, Action, 그리고 Multi-Agent.
- 현실성, 확장성, 상호작용을 기반으로 한 이차 평가 특징.
- 일곱 가지 특징에 걸친 시뮬레이터의 포괄적 질적 및 양적 비교(표 I 및 II).
- 세 가지 주요 구현형 AI 연구 과제와 최첨단 방법, 지표, 데이터 세트에 대한 고찰(표 III).
- 시뮬레이터, 데이터 세트, 과제 간의 연계 분석을 통해 도전과제를 식별한다.
실험 결과
연구 질문
- RQ1현행 구현형 AI 시뮬레이터의 현실성, 확장성, 상호작용성에 대한 가능성과 한계는 무엇인가?
- RQ2다양한 시뮬레이터가 시각 탐색, 시각 내비게이션, 구현형 QA와 같은 핵심 과제를 어떻게 지원하는가?
- RQ3연구자들이 특정 구현형 AI 과제에 적합한 시뮬레이터와 데이터 세트를 선택하는 데 도움이 되는 지침은 무엇인가?
- RQ4구현형 AI 연구 및 시뮬레이션 프레임워크의 주요 도전 과제와 향후 방향은 무엇인가?
주요 결과
- Nine embodied AI simulators (DeepMind Lab, AI2-THOR, CHALET, VirtualHome, VRKitchen, Habitat-Sim, iGibson, SAPIEN, ThreeDWorld) are benchmarked on seven features.
- Realism, scalability, and interactivity are proposed as three secondary evaluation features to compare simulators.
- AI2-THOR, iGibson, and Habitat-Sim offer broad realism, interactivity, and scalability, making them popular for diverse embodied AI tasks.
- The three main tasks—visual exploration, visual navigation, and embodied QA—cover state-of-the-art approaches, evaluation metrics, and datasets.
- The paper provides recommendations for simulator-task selections and directions for future research.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.