Skip to main content
QUICK REVIEW

[論文レビュー] Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting

Benjamin Wilson, William Qi|arXiv (Cornell University)|Jan 1, 2023
Autonomous Vehicle Technology and Safety被引用数 113
ひとこと要約

Argoverse 2 は、自己運転の認識と予測のための三つの大規模データセットを導入します:Sensor、Lidar、Motion Forecasting。豊富なHDマップと多様な都市セットを備え、3D検出、自己教師あり学習、マルチエージェント予測を推進します。

ABSTRACT

We introduce Argoverse 2 (AV2) - a collection of three datasets for perception and forecasting research in the self-driving domain. The annotated Sensor Dataset contains 1,000 sequences of multimodal data, encompassing high-resolution imagery from seven ring cameras, and two stereo cameras in addition to lidar point clouds, and 6-DOF map-aligned pose. Sequences contain 3D cuboid annotations for 26 object categories, all of which are sufficiently-sampled to support training and evaluation of 3D perception models. The Lidar Dataset contains 20,000 sequences of unlabeled lidar point clouds and map-aligned pose. This dataset is the largest ever collection of lidar sensor data and supports self-supervised learning and the emerging task of point cloud forecasting. Finally, the Motion Forecasting Dataset contains 250,000 scenarios mined for interesting and challenging interactions between the autonomous vehicle and other actors in each local scene. Models are tasked with the prediction of future motion for "scored actors" in each scenario and are provided with track histories that capture object location, heading, velocity, and category. In all three datasets, each scenario contains its own HD Map with 3D lane and crosswalk geometry - sourced from data captured in six distinct cities. We believe these datasets will support new and existing machine learning research problems in ways that existing datasets do not. All datasets are released under the CC BY-NC-SA 4.0 license.

研究の動機と目的

  • 多様で難易度の高いデータシナリオを通じて、安全で信頼性の高い自動運転を促進する。
  • HDマップと豊富な注釈を持つ三つのデータセット(Sensor, Lidar, Motion Forecasting)を提供し、認識と予測タスクを支援する。
  • 大規模で多様なデータを用いた3D物体検出、点群予測、マルチエージェント運動予測の研究を可能にする。

提案手法

  • AV2 Sensor Dataset を1,000シーン、7台のカメラ、ステレオ画像、2つのライダー、3Dキュボイドを用いて26カテゴリを30クラスにまたがって提供。
  • AV2 Lidar Dataset を20,000のラベルなしシーケンスとマップ整列ポーズを使い、自己教師あり学習と点群予測を支援。
  • AV2 Motion Forecasting Dataset を250,000のシナリオ、多様な相互作用を抽出し、各シナリオに局所マップと11秒のトラック履歴を備え、マルチアクター予測を実現。
  • 各シナリオにHDマップを提供、3D車線ジオメトリ、走行可能領域、そして高解像度の地表高を含み、マップ情報を用いた認識と予測をサポート。
  • 3D物体検出、点群予測、モーション予測のベースライン実験を提供し、データセットの有用性を示す。

実験結果

リサーチクエスチョン

  • RQ1より大きく多様な分類法と長尾の物体表現は、自己運転認識と予測モデルをどのように改善できるか。
  • RQ2HDマップとシナリオごとのマップは、認識と予測の精度と一般化にどのような利点をもたらすか。
  • RQ3大規模LiDARデータセットでの自己教師あり学習は、点群予測と表現学習にどのような影響を与えるか。
  • RQ4マップの事前知識と社会的文脈を取り入れたマルチエージェント運動予測の課題と性能傾向はどうなるか。

主な発見

  • AV2 Sensor Dataset には訓練/評価用に30の分類と26のよくサンプリングされたカテゴリを含む1,000シーンがある。
  • AV2 Lidar Dataset には10 Hzで20,000シーケンスが含まれ、自己教師あり学習と点群予測を支援する。
  • AV2 Motion Forecasting Dataset には6都市で興味深い相互作用を抽出した250,000のシナリオがあり、シナリオごとにHDマップを備える。
  • 3D物体検出ベースライン(CenterPoint)は大規模な分類とマルチヘッドアーキテクチャの有用性を、26の評価クラスに対して示している。
  • 点群予測は訓練データが増えると向上し、訓練ログが増えるにつれて平均IoU、l1ノルム、Chamfer距離で安定した改善を示す。
  • モーション予測のベースラインはマップ事前知識と社会的文脈の重要性を示し、グラフベースのアテンション手法はより長い予測区間で強力な性能を発揮。
  • 基準となるモーション予測の結果は Argoverse 1.1 より難易度が増していることを示し、このデータセットの長尾と多様モードの課題を浮き彫りにしている。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。