[Paper Review] Computer Vision for Autonomous Vehicles: Problems, Datasets and State of the Art
A comprehensive survey of perception problems, datasets, and state-of-the-art methods for autonomous driving, with benchmark analyses and an accompanying online resource.
Recent years have witnessed enormous progress in AI-related fields such as computer vision, machine learning, and autonomous vehicles. As with any rapidly growing field, it becomes increasingly difficult to stay up-to-date or enter the field as a beginner. While several survey papers on particular sub-problems have appeared, no comprehensive survey on problems, datasets, and methods in computer vision for autonomous vehicles has been published. This book attempts to narrow this gap by providing a survey on the state-of-the-art datasets and techniques. Our survey includes both the historically most relevant literature as well as the current state of the art on several specific topics, including recognition, reconstruction, motion estimation, tracking, scene understanding, and end-to-end learning for autonomous driving. Towards this goal, we analyze the performance of the state of the art on several challenging benchmarking datasets, including KITTI, MOT, and Cityscapes. Besides, we discuss open problems and current research challenges. To ease accessibility and accommodate missing references, we also provide a website that allows navigating topics as well as methods and provides additional information.
Motivation & Objective
- Survey the history and key challenges in autonomous-vision perception for driving.
- Summarize major datasets, benchmarks, and evaluation metrics across perception tasks.
- Review state-of-the-art methods for detection, tracking, segmentation, reconstruction, motion estimation, and scene understanding.
- Discuss open problems, research challenges, and directions for end-to-end learning in autonomous driving.
- Provide accessible navigation of topics and methods via an online tool.
Proposed method
- Systematic review of perception-related modules in modular autonomous driving pipelines and end-to-end approaches.
- Comparative analysis of state-of-the-art techniques on popular datasets (e.g., KITTI, MOT, Cityscapes).
- Discussion of sensor suites, camera models, and calibration relevant to autonomous vision.
- Overview of datasets and benchmarks, including synthetic data generation, with qualitative and quantitative insights.
- Provision of an interactive online resource visualizing surveyed papers and methods.
Experimental results
Research questions
- RQ1What are the main perception tasks, datasets, and benchmarks used in autonomous-vehicle vision research?
- RQ2What are the current state-of-the-art methods for detection, tracking, segmentation, reconstruction, and motion estimation in driving scenarios?
- RQ3How do different datasets and benchmarks compare in realism, diversity, and evaluation?
- RQ4What open problems and challenges remain for perception in autonomous driving and for end-to-end learning approaches?
- RQ5How can researchers efficiently navigate the field and access summarized information and relationships among works?
Key findings
- The paper surveys perception-related modules and end-to-end approaches across recognition, reconstruction, motion estimation, tracking, scene understanding, and end-to-end driving.
- It analyzes state-of-the-art techniques on challenging benchmarks such as KITTI, MOT, and Cityscapes.
- It discusses sensor suites, camera models, and calibration for robust autonomous-vision systems.
- An online interactive tool is provided to navigate topics, methods, and references.
- The work highlights open problems and research challenges in autonomous-vision perception and end-to-end driving.
- Synthetic data generation and its role in benchmarking and training are reviewed as part of dataset discussions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.