[论文解读] Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting
Argoverse 2 引入面向自驾感知与预测的三大大规模数据集:Sensor、Lidar 和 Motion Forecasting,配有丰富的 HD 地图和多样的城市集合,以推动 3D 检测、自监督学习与多智能体预测。
We introduce Argoverse 2 (AV2) - a collection of three datasets for perception and forecasting research in the self-driving domain. The annotated Sensor Dataset contains 1,000 sequences of multimodal data, encompassing high-resolution imagery from seven ring cameras, and two stereo cameras in addition to lidar point clouds, and 6-DOF map-aligned pose. Sequences contain 3D cuboid annotations for 26 object categories, all of which are sufficiently-sampled to support training and evaluation of 3D perception models. The Lidar Dataset contains 20,000 sequences of unlabeled lidar point clouds and map-aligned pose. This dataset is the largest ever collection of lidar sensor data and supports self-supervised learning and the emerging task of point cloud forecasting. Finally, the Motion Forecasting Dataset contains 250,000 scenarios mined for interesting and challenging interactions between the autonomous vehicle and other actors in each local scene. Models are tasked with the prediction of future motion for "scored actors" in each scenario and are provided with track histories that capture object location, heading, velocity, and category. In all three datasets, each scenario contains its own HD Map with 3D lane and crosswalk geometry - sourced from data captured in six distinct cities. We believe these datasets will support new and existing machine learning research problems in ways that existing datasets do not. All datasets are released under the CC BY-NC-SA 4.0 license.
研究动机与目标
- 通过多样且具挑战性的数据场景,推动安全、可靠的自动驾驶。
- 提供三个数据集(Sensor、Lidar、Motion Forecasting),具备HD地图和丰富注释,用于感知与预测任务。
- 利用大规模、多样化数据,开展3D目标检测、点云预测和多智能体运动预测研究。
提出的方法
- 引入 AV2 Sensor Dataset,具有 1,000 个场景、7 个相机、立体影像、2 个激光雷达、覆盖 30 个类中的 26 个类的 3D 矩形框。
- 引入 AV2 Lidar Dataset,包含 20,000 个未标注序列和与地图对齐的位姿,用于自监督学习和点云预测。
- 引入 AV2 Motion Forecasting Dataset,包含 250,000 个用于多样交互的场景,每个场景配有局部地图和 11 s 的轨迹历史,用于多智能体预测。
- 为每个场景提供 HD 地图,包括 3D 车道几何、可行驶区域和高分辨率地面高程,以支持基于地图的信息感知与预测。
- 提供 3D 目标检测、点云预测和运动预测的基线实验,以展示数据集的实用性。
实验结果
研究问题
- RQ1更大规模、多样化的分类体系和长尾对象表示如何提升自动驾驶感知与预测模型?
- RQ2HD 地图和每个场景地图对感知与预测的准确性与泛化能力有何帮助?
- RQ3对大型 LiDAR 数据集进行自监督学习如何影响点云预测与表示学习?
- RQ4在引入地图先验与社会上下文时,多智能体运动预测面临哪些挑战与性能趋势?
主要发现
- AV2 Sensor Dataset 包含 1,000 个场景,具有 30 个分类学与 26 个样本充足的类别用于训练/评估。
- AV2 Lidar Dataset 包含 以 10 Hz 采样的 20,000 条序列,支持自监督学习与点云预测。
- AV2 Motion Forecasting Dataset 包含在六个城市中挖掘出用于有趣交互的 250,000 个场景,且每个场景具备 HD 地图。
- 3D 目标检测基线(CenterPoint)证明了大分类体系和多头结构在 26 个评估类中的实用性。
- 随着训练数据增多,点云预测得到提升,平均 IoU、l1 范数和 Chamfer 距离在训练日志增加时呈现持续收益。
- 运动预测基线显示地图先验与社会上下文的重要性,图注意力方法在较高预测 horizon 上表现出色。
- 基线运动预测结果表明相较于 Argverse 1.1 的难度有所增加,突出数据集的长尾和多模态挑战。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。