[Paper Review] NuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles
Introduces the world’s first closed-loop ML-based planning benchmark for autonomous driving, with a large real-world dataset, a lightweight closed-loop simulator, and planning-specific metrics.
In this work, we propose the world's first closed-loop ML-based planning benchmark for autonomous driving. While there is a growing body of ML-based motion planners, the lack of established datasets and metrics has limited the progress in this area. Existing benchmarks for autonomous vehicle motion prediction have focused on short-term motion forecasting, rather than long-term planning. This has led previous works to use open-loop evaluation with L2-based metrics, which are not suitable for fairly evaluating long-term planning. Our benchmark overcomes these limitations by introducing a large-scale driving dataset, lightweight closed-loop simulator, and motion-planning-specific metrics. We provide a high-quality dataset with 1500h of human driving data from 4 cities across the US and Asia with widely varying traffic patterns (Boston, Pittsburgh, Las Vegas and Singapore). We will provide a closed-loop simulation framework with reactive agents and provide a large set of both general and scenario-specific planning metrics. We plan to release the dataset at NeurIPS 2021 and organize benchmark challenges starting in early 2022.
Motivation & Objective
- Provide a large-scale, real-world dataset for ML-based planning in autonomous driving.
- Introduce a closed-loop planning benchmark to evaluate interactions with other agents.
- Define planning-specific metrics including safety, similarity to human driving, and scenario-based assessments.
- Enable community-driven development and standardization of ML-based planning evaluations.
Proposed method
- Release 1500 hours of autolabeled driving data from four cities (Las Vegas, Boston, Pittsburgh, Singapore) with sensor suites and semantic maps.
- Provide a lightweight closed-loop simulator with reactive agent modeling and a controller that follows planned trajectories.
- Define planning-centric evaluation metrics spanning traffic rule compliance, human driving similarity, vehicle dynamics, goal achievement, and scenario-based cases.
- Support open-loop and closed-loop evaluation, including reactive and non-reactive closed-loop tasks, through a containerized evaluation framework.
- Annotate scenarios automatically with tags (e.g., merges, lane changes, protected turns) to enable scenario-based metrics.
Experimental results
Research questions
- RQ1How can a closed-loop ML-based planning benchmark be formed from real-world driving data?
- RQ2What metrics best capture planning quality, safety, and goal achievement in autonomous driving?
- RQ3How do ML-based planners perform under non-reactive vs. reactive closed-loop conditions?
- RQ4Can a standardized benchmark accelerate progress in ML-based planning across geographies and traffic patterns?
Key findings
- Proposes the first public benchmark for real-world data with closed-loop planner evaluation.
- Offers the largest existing public real-world AV dataset with high-quality autolabeled tracks from 4 cities.
- Lists a comprehensive suite of planning-specific metrics, including traffic rule violations, human driving similarity, vehicle dynamics, and goal achievement.
- Introduces an open-grounded evaluation workflow with containerized planner submissions to enable reproducible closed-loop testing.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.