Skip to main content
QUICK REVIEW

[Paper Review] Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting

Benjamin Wilson, William Qi|arXiv (Cornell University)|Jan 1, 2023
Autonomous Vehicle Technology and Safety113 citations
TL;DR

Argoverse 2 introduces three large-scale datasets for perception and forecasting in self-driving: Sensor, Lidar, and Motion Forecasting, with rich HD maps and a diverse city set to advance 3D detection, self-supervised learning, and multi-agent forecasting.

ABSTRACT

We introduce Argoverse 2 (AV2) - a collection of three datasets for perception and forecasting research in the self-driving domain. The annotated Sensor Dataset contains 1,000 sequences of multimodal data, encompassing high-resolution imagery from seven ring cameras, and two stereo cameras in addition to lidar point clouds, and 6-DOF map-aligned pose. Sequences contain 3D cuboid annotations for 26 object categories, all of which are sufficiently-sampled to support training and evaluation of 3D perception models. The Lidar Dataset contains 20,000 sequences of unlabeled lidar point clouds and map-aligned pose. This dataset is the largest ever collection of lidar sensor data and supports self-supervised learning and the emerging task of point cloud forecasting. Finally, the Motion Forecasting Dataset contains 250,000 scenarios mined for interesting and challenging interactions between the autonomous vehicle and other actors in each local scene. Models are tasked with the prediction of future motion for "scored actors" in each scenario and are provided with track histories that capture object location, heading, velocity, and category. In all three datasets, each scenario contains its own HD Map with 3D lane and crosswalk geometry - sourced from data captured in six distinct cities. We believe these datasets will support new and existing machine learning research problems in ways that existing datasets do not. All datasets are released under the CC BY-NC-SA 4.0 license.

Motivation & Objective

  • Motivate safe, reliable autonomous driving through diverse, challenging data scenarios.
  • Provide three datasets (Sensor, Lidar, Motion Forecasting) with HD maps and rich annotations for perception and forecasting tasks.
  • Enable studies in 3D object detection, point cloud forecasting, and multi-agent motion forecasting using large-scale, diverse data.

Proposed method

  • Introduce AV2 Sensor Dataset with 1,000 scenes, 7 cameras, stereo imagery, 2 lidars, 3D cuboids for 26 categories across 30 classes.
  • Introduce AV2 Lidar Dataset with 20,000 unlabeled sequences and map-aligned poses for self-supervised learning and point cloud forecasting.
  • Introduce AV2 Motion Forecasting Dataset with 250,000 scenarios mined for diverse interactions, each with local maps and 11 s of track history for multi-actor forecasting.
  • Provide HD maps per scenario including 3D lane geometry, driveable areas, and high-resolution ground height to support map-informed perception and forecasting.
  • Offer baseline experiments for 3D object detection, point cloud forecasting, and motion forecasting to demonstrate dataset utility.

Experimental results

Research questions

  • RQ1How can larger, diverse taxonomies and long-tail object representations improve self-driving perception and forecasting models?
  • RQ2What benefits do HD maps and per-scenario maps provide for perception and forecasting accuracy and generalization?
  • RQ3How does self-supervised learning on large LiDAR datasets impact point cloud forecasting and representation learning?
  • RQ4What are the challenges and performance trends in multi-agent motion forecasting when incorporating map priors and social context?

Key findings

  • AV2 Sensor Dataset contains 1,000 scenes with 30 taxonomic classes and 26 well-sampled categories for training/evaluation.
  • AV2 Lidar Dataset contains 20,000 sequences at 10 Hz, supporting self-supervised learning and point cloud forecasting.
  • AV2 Motion Forecasting Dataset contains 250,000 scenarios mined for interesting interactions across six cities with per-scenario HD maps.
  • 3D object detection baseline (CenterPoint) demonstrates the utility of the large taxonomy and multi-head architecture for the 26 evaluated classes.
  • Point cloud forecasting improves with more training data, showing steady gains in mean IoU, l1-norm, and Chamfer distance as training logs increase.
  • Motion forecasting baselines reveal the importance of map priors and social context, with graph-based attention methods providing strong performance at higher prediction horizons.
  • Baseline motion forecasting results indicate increased difficulty compared to Argoverse 1.1, highlighting the dataset’s long-tail and multimodal challenges.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.