[Paper Review] Deep Learning for Visual Localization and Mapping: A Survey
This survey provides a comprehensive taxonomy and analysis of deep learning-based visual localization and mapping methods, covering visual odometry, global relocalization, mapping, and SLAM. It evaluates the promise of deep learning in improving robustness and accuracy over traditional methods, while identifying key challenges in deployment, scalability, interpretability, and efficiency on resource-constrained systems.
Deep learning based localization and mapping approaches have recently emerged as a new research direction and receive significant attentions from both industry and academia. Instead of creating hand-designed algorithms based on physical models or geometric theories, deep learning solutions provide an alternative to solve the problem in a data-driven way. Benefiting from the ever-increasing volumes of data and computational power on devices, these learning methods are fast evolving into a new area that shows potentials to track self-motion and estimate environmental model accurately and robustly for mobile agents. In this work, we provide a comprehensive survey, and propose a taxonomy for the localization and mapping methods using deep learning. This survey aims to discuss two basic questions: whether deep learning is promising to localization and mapping; how deep learning should be applied to solve this problem. To this end, a series of localization and mapping topics are investigated, from the learning based visual odometry, global relocalization, to mapping, and simultaneous localization and mapping (SLAM). It is our hope that this survey organically weaves together the recent works in this vein from robotics, computer vision and machine learning communities, and serves as a guideline for future researchers to apply deep learning to tackle the problem of visual localization and mapping.
Motivation & Objective
- To evaluate whether deep learning is a promising approach for visual localization and mapping compared to traditional model-based methods.
- To identify and categorize key deep learning techniques applied to visual odometry, global relocalization, mapping, and SLAM.
- To analyze the limitations of current deep learning-based approaches, including generalization, interpretability, and computational cost.
- To guide future research by highlighting open challenges such as real-world deployment, scalability, safety, and model efficiency on edge devices.
Proposed method
- Proposes a structured taxonomy for deep learning-based visual localization and mapping, organizing methods by task: visual odometry, global relocalization, mapping, and SLAM.
- Reviews data-driven approaches that learn feature representations and implicit neural mappings from large-scale datasets, replacing handcrafted geometric models.
- Examines self-supervised learning techniques using novel view synthesis to generate supervisory signals for pose and depth estimation from unlabeled video.
- Integrates learned depth estimates to resolve scale ambiguity in monocular SLAM, improving absolute pose accuracy.
- Explores uncertainty estimation and confidence metrics to enhance reliability and safety in critical applications.
- Analyzes trade-offs between model accuracy, size, inference speed, and hardware constraints, advocating for efficient network design in edge deployment.

Experimental results
Research questions
- RQ1To what extent can deep learning improve the accuracy and robustness of visual localization and mapping compared to classical geometric methods?
- RQ2How can self-supervised learning and neural implicit representations be leveraged to reduce reliance on annotated data in visual SLAM?
- RQ3What are the key limitations of deep learning-based approaches in real-world deployment, particularly regarding computational cost, model size, and energy efficiency?
- RQ4How can interpretability and uncertainty quantification be integrated into deep learning models to enhance safety and reliability in autonomous systems?
- RQ5What trade-offs exist between model performance, generalization, and efficiency, and how can they be optimized for deployment on mobile and wearable platforms?
Key findings
- Deep learning-based methods significantly improve robustness and accuracy in challenging conditions such as poor lighting and dynamic scenes, outperforming traditional geometric approaches.
- Self-supervised learning using novel view synthesis enables end-to-end training of SLAM systems with no need for ground-truth poses or depth, achieving competitive performance on benchmark datasets.
- Leveraging learned depth estimates resolves the scale-ambiguity problem in monocular SLAM, enabling absolute pose estimation without external sensors.
- Uncertainty estimation in deep models provides a confidence metric that can flag unreliable predictions, enhancing system safety in critical applications.
- Despite high accuracy, deep learning models often suffer from poor generalization to out-of-domain environments and require substantial computational resources, especially on edge devices.
- Current approaches remain limited in scalability, with most methods tested only on urban or indoor scenes, and few capable of large-scale, complex environment reconstruction.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.