Skip to main content
QUICK REVIEW

[Paper Review] AeroVerse: UAV-Agent Benchmark Suite for Simulating, Pre-training, Finetuning, and Evaluating Aerospace Embodied World Models

Fanglong Yao, Yuanchang Yue|arXiv (Cornell University)|Aug 28, 2024
Simulation Techniques and ApplicationsDecision Sciences3 citations
TL;DR

This paper introduces AeroVerse, a comprehensive benchmark suite for aerospace embodied world models in UAVs, integrating simulation, pre-training, fine-tuning, and evaluation. It proposes the first large-scale real and virtual image-text datasets (AerialAgent-Ego10k and CyberAgent-Ego500k), defines five novel downstream tasks, and introduces evaluation via GPT-4-based SkyAgent-Eval, enabling end-to-end autonomous UAV intelligence with demonstrated performance gains across 2D/3D visual-language models.

ABSTRACT

Aerospace embodied intelligence aims to empower unmanned aerial vehicles (UAVs) and other aerospace platforms to achieve autonomous perception, cognition, and action, as well as egocentric active interaction with humans and the environment. The aerospace embodied world model serves as an effective means to realize the autonomous intelligence of UAVs and represents a necessary pathway toward aerospace embodied intelligence. However, existing embodied world models primarily focus on ground-level intelligent agents in indoor scenarios, while research on UAV intelligent agents remains unexplored. To address this gap, we construct the first large-scale real-world image-text pre-training dataset, AerialAgent-Ego10k, featuring urban drones from a first-person perspective. We also create a virtual image-text-pose alignment dataset, CyberAgent Ego500k, to facilitate the pre-training of the aerospace embodied world model. For the first time, we clearly define 5 downstream tasks, i.e., aerospace embodied scene awareness, spatial reasoning, navigational exploration, task planning, and motion decision, and construct corresponding instruction datasets, i.e., SkyAgent-Scene3k, SkyAgent-Reason3k, SkyAgent-Nav3k and SkyAgent-Plan3k, and SkyAgent-Act3k, for fine-tuning the aerospace embodiment world model. Simultaneously, we develop SkyAgentEval, the downstream task evaluation metrics based on GPT-4, to comprehensively, flexibly, and objectively assess the results, revealing the potential and limitations of 2D/3D visual language models in UAV-agent tasks. Furthermore, we integrate over 10 2D/3D visual-language models, 2 pre-training datasets, 5 finetuning datasets, more than 10 evaluation metrics, and a simulator into the benchmark suite, i.e., AeroVerse, which will be released to the community to promote exploration and development of aerospace embodied intelligence.

Motivation & Objective

  • To address the lack of comprehensive benchmarks for UAV-based embodied intelligence in aerospace applications.
  • To close the research gap in UAV embodied world models by developing a unified framework for simulation, pre-training, fine-tuning, and evaluation.
  • To define and structure five novel downstream tasks—scene awareness, spatial reasoning, navigation, task planning, and motion decision—for UAV agents.
  • To create high-fidelity, first-person perspective datasets (AerialAgent-Ego10k and CyberAgent-Ego500k) to support pre-training of visual-language models for aerial agents.
  • To develop SkyAgent-Eval, a GPT-4-powered evaluation suite that enables objective, flexible, and comprehensive assessment of model performance across diverse UAV tasks.

Proposed method

  • Developed AeroSimulator, a simulation platform with four realistic urban scenes for UAV flight simulation under dynamic and observable environmental conditions.
  • Constructed AerialAgent-Ego10k, a large-scale real-world image-text dataset using first-person drone footage to support pre-training of aerial embodied world models.
  • Created CyberAgent-Ego500k, a synthetic virtual dataset with aligned image-text-pose annotations to enhance pre-training generalization and data efficiency.
  • Defined five downstream tasks—scene awareness, spatial reasoning, navigational exploration, task planning, and motion decision—each with dedicated instruction-tuning datasets (SkyAgent-Scene3k to SkyAgent-Act3k).
  • Designed SkyAgent-Eval, a multi-metric evaluation framework based on GPT-4 that assesses model outputs across multiple dimensions, including accuracy, coherence, and task-specific reasoning.
  • Integrated over 10 2D/3D visual-language models, two pre-training datasets, five fine-tuning datasets, and more than 10 evaluation metrics into the unified AeroVerse benchmark suite.

Experimental results

Research questions

  • RQ1How can a comprehensive benchmark suite be designed to support the full pipeline of UAV-agent development, from simulation to evaluation?
  • RQ2What are the key challenges in defining and structuring downstream tasks for UAV-based embodied intelligence, particularly in 4D space-time and partial observability?
  • RQ3To what extent do 2D and 3D visual-language models generalize across diverse urban scenes and task types in UAV environments?
  • RQ4How effective is GPT-4-based evaluation (SkyAgent-Eval) in objectively measuring and revealing the strengths and limitations of visual-language models in aerospace embodied tasks?
  • RQ5What impact do model size and architecture have on performance in UAV-specific embodied reasoning tasks?

Key findings

  • The Qwen-LV-7B model achieved the highest average BLEU scores across all four urban scenes, demonstrating strong generalization and robustness in aerial scene understanding.
  • GPT-4o and GPT-4-vision-review outperformed other models in trajectory description tasks, delivering the most detailed and timeline-accurate flight path interpretations.
  • BLIP2-flan-t5-xxl exhibited poor adherence to task-specific instructions, often generating image caption-like outputs instead of structured flight path descriptions.
  • Models such as InstructBLIP and BLIP2 showed superior performance in scene captioning (Task 1), while LLaVA and Mplug series models excelled in multi-modal reasoning (Task 4).
  • Increasing model parameters from 7B to 13B did not consistently improve performance, indicating that scaling alone does not guarantee better generalization in UAV embodied tasks.
  • The 3D-LLM model underperformed significantly due to challenges in processing 3D scene inputs compared to 2D image-based representations, highlighting architectural limitations in 3D reasoning.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.