Skip to main content
QUICK REVIEW

[Paper Review] Vision Meets Drones: Past, Present and Future

Pengfei Zhu, Longyin Wen|arXiv (Cornell University)|Jan 16, 2020
Video Surveillance and Tracking MethodsComputer Science172 references140 citations
TL;DR

This paper introduces VisDrone, a large-scale, fully annotated drone-captured dataset comprising four tracks—image detection, video detection, single-object tracking, and multi-object tracking—collected across 14 Chinese cities. It provides a benchmark for evaluating and advancing visual analysis algorithms on drones, significantly boosting research in aerial video understanding through extensive evaluation and future research directions.

ABSTRACT

Drones, or general UAVs, equipped with cameras have been fast deployed with a wide range of applications, including agriculture, aerial photography, and surveillance. Consequently, automatic understanding of visual data collected from drones becomes highly demanding, bringing computer vision and drones more and more closely. To promote and track the evelopments of object detection and tracking algorithms, we have organized two challenge workshops in conjunction with ECCV 2018, and ICCV 2019, attracting more than 100 teams around the world. We provide a large-scale drone captured dataset, VisDrone, which includes four tracks, i.e., (1) image object detection, (2) video object detection, (3) single object tracking, and (4) multi-object tracking. In this paper, we first presents a thorough review of object detection and tracking datasets and benchmarks, and discuss the challenges of collecting large-scale drone-based object detection and tracking datasets with fully manual annotations. After that, we describe our VisDrone dataset, which is captured over various urban/suburban areas of 14 different cities across China from North to South. Being the largest such dataset ever published, VisDrone enables extensive evaluation and investigation of visual analysis algorithms on the drone platform. We provide a detailed analysis of the current state of the field of large-scale object detection and tracking on drones, and conclude the challenge as well as propose future directions. We expect the benchmark largely boost the research and development in video analysis on drone platforms. All the datasets and experimental results can be downloaded from the website: this https URL.

Motivation & Objective

  • To address the growing demand for automatic visual understanding of drone-captured data in applications like agriculture, surveillance, and aerial photography.
  • To overcome challenges in collecting large-scale, fully annotated drone datasets with consistent quality and diversity.
  • To establish a benchmark for evaluating object detection and tracking algorithms on drone platforms through a comprehensive dataset and challenge workshops.
  • To promote research progress by organizing international challenges at ECCV 2018 and ICCV 2019, attracting over 100 global teams.
  • To provide a foundation for future advancements in drone-based video analysis through detailed analysis and open access to data and results.

Proposed method

  • The VisDrone dataset was collected from diverse urban and suburban areas across 14 cities in China, spanning from north to south to ensure geographical and environmental diversity.
  • The dataset includes four distinct tracks: image object detection, video object detection, single object tracking, and multi-object tracking, each with fully manual annotations.
  • The dataset is the largest publicly available drone-based benchmark for visual analysis, supporting extensive evaluation of algorithms.
  • The authors organized two international challenge workshops at ECCV 2018 and ICCV 2019 to evaluate and track algorithmic progress on the VisDrone dataset.
  • All data and results are publicly available via a dedicated website to encourage open research and reproducibility.
  • The paper provides a comprehensive review of existing datasets and benchmarks in drone-based object detection and tracking, identifying key limitations and opportunities.

Experimental results

Research questions

  • RQ1What are the key challenges in collecting large-scale, fully annotated drone-based datasets for object detection and tracking?
  • RQ2How does the VisDrone dataset compare to existing benchmarks in terms of scale, diversity, and annotation quality?
  • RQ3What are the current performance limits and bottlenecks in drone-based object detection and tracking algorithms?
  • RQ4How can large-scale benchmarks like VisDrone accelerate progress in aerial video analysis?
  • RQ5What future research directions are most promising for advancing visual analysis on drone platforms?

Key findings

  • VisDrone is the largest publicly available drone-captured dataset for object detection and tracking, collected across 14 diverse cities in China.
  • The dataset supports four distinct tasks: image detection, video detection, single-object tracking, and multi-object tracking, each with fully manual annotations.
  • The dataset has enabled international benchmarking through challenges at ECCV 2018 and ICCV 2019, attracting over 100 teams worldwide.
  • The authors identify significant challenges in data collection, including annotation consistency, scale, and environmental variability across regions.
  • The paper concludes that VisDrone provides a robust foundation for advancing research in drone-based video analysis and proposes future directions for algorithmic development.
  • All dataset and results are publicly accessible via a dedicated website, promoting open science and reproducibility in the field.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.