[论文解读] AeroVerse: UAV-Agent Benchmark Suite for Simulating, Pre-training, Finetuning, and Evaluating Aerospace Embodied World Models
本文提出了AeroVerse,一个面向无人机(UAV)航空航天具身世界模型的综合性基准测试套件,整合了仿真、预训练、微调和评估。该工作构建了首个大规模真实与虚拟图像-文本数据集(AerialAgent-Ego10k 和 CyberAgent-Ego500k),定义了五个新型下游任务,并引入基于GPT-4的SkyAgent-Eval评估方法,实现了从2D/3D视觉-语言模型到端到端自主无人机智能的性能提升。
Aerospace embodied intelligence aims to empower unmanned aerial vehicles (UAVs) and other aerospace platforms to achieve autonomous perception, cognition, and action, as well as egocentric active interaction with humans and the environment. The aerospace embodied world model serves as an effective means to realize the autonomous intelligence of UAVs and represents a necessary pathway toward aerospace embodied intelligence. However, existing embodied world models primarily focus on ground-level intelligent agents in indoor scenarios, while research on UAV intelligent agents remains unexplored. To address this gap, we construct the first large-scale real-world image-text pre-training dataset, AerialAgent-Ego10k, featuring urban drones from a first-person perspective. We also create a virtual image-text-pose alignment dataset, CyberAgent Ego500k, to facilitate the pre-training of the aerospace embodied world model. For the first time, we clearly define 5 downstream tasks, i.e., aerospace embodied scene awareness, spatial reasoning, navigational exploration, task planning, and motion decision, and construct corresponding instruction datasets, i.e., SkyAgent-Scene3k, SkyAgent-Reason3k, SkyAgent-Nav3k and SkyAgent-Plan3k, and SkyAgent-Act3k, for fine-tuning the aerospace embodiment world model. Simultaneously, we develop SkyAgentEval, the downstream task evaluation metrics based on GPT-4, to comprehensively, flexibly, and objectively assess the results, revealing the potential and limitations of 2D/3D visual language models in UAV-agent tasks. Furthermore, we integrate over 10 2D/3D visual-language models, 2 pre-training datasets, 5 finetuning datasets, more than 10 evaluation metrics, and a simulator into the benchmark suite, i.e., AeroVerse, which will be released to the community to promote exploration and development of aerospace embodied intelligence.
研究动机与目标
- 为解决航空航天应用中基于无人机的具身智能缺乏综合性基准的问题。
- 通过构建统一的仿真、预训练、微调与评估框架,弥合无人机具身世界模型的研究空白。
- 定义并组织五个新型下游任务——场景感知、空间推理、导航、任务规划与运动决策,用于无人机智能体。
- 创建高保真度、第一人称视角的数据集(AerialAgent-Ego10k 与 CyberAgent-Ego500k),以支持空中智能体视觉-语言模型的预训练。
- 开发SkyAgent-Eval,一个基于GPT-4的评估套件,实现对多样化无人机任务中模型性能的客观、灵活且全面的评估。
提出的方法
- 开发AeroSimulator,一个包含四个真实城市场景的仿真平台,用于在动态且可观测的环境条件下进行无人机飞行仿真。
- 构建AerialAgent-Ego10k,一个大规模真实世界图像-文本数据集,利用第一人称无人机影像支持空中具身世界模型的预训练。
- 创建CyberAgent-Ego500k,一个带有对齐图像-文本-位姿标注的合成虚拟数据集,以增强预训练的泛化能力与数据效率。
- 定义五个下游任务——场景感知、空间推理、导航探索、任务规划与运动决策,每个任务均配备专用指令微调数据集(SkyAgent-Scene3k 至 SkyAgent-Act3k)。
- 设计SkyAgent-Eval,一个基于GPT-4的多指标评估框架,从准确性、连贯性及任务特定推理等多个维度评估模型输出。
- 将超过10个2D/3D视觉-语言模型、两个预训练数据集、五个微调数据集以及十余项评估指标整合进统一的AeroVerse基准测试套件中。
实验结果
研究问题
- RQ1如何设计一个综合性基准测试套件,以支持从仿真到评估的无人机智能体开发全流程?
- RQ2在4D时空与部分可观测性条件下,定义和构建无人机具身智能下游任务的关键挑战是什么?
- RQ32D与3D视觉-语言模型在多样化城市场景与任务类型中的泛化能力如何,特别是在无人机环境中?
- RQ4基于GPT-4的评估(SkyAgent-Eval)在客观衡量并揭示视觉-语言模型在航空航天具身任务中优势与局限性方面的有效性如何?
- RQ5模型参数量与架构对无人机特定具身推理任务性能的影响是什么?
主要发现
- Qwen-LV-7B模型在所有四个城市场景中均取得最高平均BLEU分数,展现出在空中场景理解方面强大的泛化能力与鲁棒性。
- GPT-4o与GPT-4-vision-review在轨迹描述任务中表现最佳,提供了最详细且时间线准确的飞行路径解读。
- BLIP2-flan-t5-xxl在任务指令遵循方面表现较差,常生成类似图像字幕的输出,而非结构化的飞行路径描述。
- InstructBLIP与BLIP2在场景字幕生成(任务1)中表现优异,而LLaVA与Mplug系列模型在多模态推理(任务4)中表现更优。
- 将模型参数从7B增至13B并未持续提升性能,表明仅靠参数量扩展无法保证在无人机具身任务中实现更好泛化。
- 3D-LLM模型表现显著较差,原因在于其处理3D场景输入的难度远高于2D图像表示,凸显了当前3D推理在架构上的局限性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。