[Paper Review] MmWave Radar and Vision Fusion for Object Detection in Autonomous Driving: A Review
This paper reviews mmWave radar and vision fusion techniques for 3D object detection in autonomous driving, categorizing fusion methods into data-level, feature-level, and decision-level approaches. It highlights the complementary strengths of radar (robust in adverse weather) and vision (rich texture and shape details), and discusses emerging trends like lidar-vision fusion and multimodal sensing for improved detection accuracy and robustness.
With autonomous driving developing in a booming stage, accurate object detection in complex scenarios attract wide attention to ensure the safety of autonomous driving. Millimeter wave (mmWave) radar and vision fusion is a mainstream solution for accurate obstacle detection. This article presents a detailed survey on mmWave radar and vision fusion based obstacle detection methods. First, we introduce the tasks, evaluation criteria, and datasets of object detection for autonomous driving. The process of mmWave radar and vision fusion is then divided into three parts: sensor deployment, sensor calibration, and sensor fusion, which are reviewed comprehensively. Specifically, we classify the fusion methods into data level, decision level, and feature level fusion methods. In addition, we introduce three-dimensional(3D) object detection, the fusion of lidar and vision in autonomous driving and multimodal information fusion, which are promising for the future. Finally, we summarize this article.
Motivation & Objective
- To provide a comprehensive survey of mmWave radar and vision fusion methods for object detection in autonomous driving.
- To analyze the challenges in complex driving scenarios, including adverse weather, occlusion, and variable object scales.
- To classify and compare fusion strategies across data-level, feature-level, and decision-level integration.
- To examine the role of 3D object detection and the integration of lidar with vision and radar for improved perception.
- To explore multimodal information fusion as a promising future direction for robust autonomous perception.
Proposed method
- The paper conducts a systematic review of existing literature on mmWave radar and vision fusion, focusing on sensor deployment, calibration, and fusion strategies.
- It classifies fusion techniques into three levels: data-level (early fusion of raw signals), feature-level (concatenation of deep features), and decision-level (fusion of detection results).
- The review includes analysis of 3D object detection methods that leverage radar point clouds and camera features to estimate bounding boxes in 3D space.
- It examines lidar-vision fusion techniques, such as using lidar to generate object proposals for image-based detection networks.
- The paper discusses multimodal fusion by treating raw radar, vision, and lidar data as distinct sensory modalities to be jointly processed.
- It evaluates performance using standard metrics like mAP, IoU, and MaxF score, and references benchmark datasets such as KITTI.
Experimental results
Research questions
- RQ1How do mmWave radar and vision complement each other in object detection under adverse weather and occlusion conditions?
- RQ2What are the key differences and trade-offs between data-level, feature-level, and decision-level fusion in radar-vision systems?
- RQ3How can 3D object detection be improved by fusing mmWave radar and vision data?
- RQ4What role does lidar-vision fusion play in enhancing detection accuracy and robustness in autonomous driving?
- RQ5What are the challenges and opportunities in multimodal information fusion for autonomous perception systems?
Key findings
- mmWave radar maintains reliable detection performance in adverse weather such as fog, rain, and snow, where vision sensors degrade significantly.
- Feature-level fusion methods, such as those using spatial attention fusion (SAF), have shown improved detection accuracy by effectively aligning radar and visual features.
- Lidar-vision fusion methods, such as those using point cloud-based proposals, achieved a MaxF score of 96.03% on the KITTI dataset for road segmentation.
- Data-level fusion of radar and vision enables earlier integration of raw sensor signals, preserving more information and improving detection consistency.
- Multimodal fusion of original radar data (not post-processed) with vision can preserve more environmental information than using processed radar outputs.
- Deep learning-based fusion networks, including those using DCNNs and fully convolutional networks (FCNs), outperform traditional hand-crafted feature methods in object classification and detection tasks.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.