Home / Journals / CMC / Online First / doi:10.32604/cmc.2026.086341
Special Issues
Table of Content

Open Access

REVIEW

Fusion-Oriented Deep Learning-Enhanced Visual SLAM: A Review

Xiruo Chen, Qi Ouyang*, Sihong Meng, Yuke Meng
Department of Automation, Chongqing University, Chongqing, China
* Corresponding Author: Qi Ouyang. Email: email

Computers, Materials & Continua https://doi.org/10.32604/cmc.2026.086341

Received 28 May 2026; Accepted 03 July 2026; Published online 28 July 2026

Abstract

Visual simultaneous localization and mapping (VSLAM) is a key technology for mobile robotics, autonomous driving, and embodied intelligence, enabling self-localization, environment reconstruction, and scene understanding. Although conventional geometric methods have achieved notable success, their performance often degrades in challenging conditions, such as low-texture scenes, severe illumination changes, dynamic interference, and long-term environmental variations. Recent advances in deep learning have created new opportunities to improve VSLAM through stronger feature representations, learned priors, semantic perception, and emerging map representations. At the same time, the increasing adoption of learning-based modules has raised important questions about integration strategies, generalization, interpretability, and real-time deployment. This paper presents a systematic review of deep learning-enhanced VSLAM, with a particular focus on how learning models are incorporated into classical simultaneous localization and mapping (SLAM) pipelines and how they function within the overall system. To provide a unified perspective, existing methods are organized into five categories according to their fusion interfaces with geometric SLAM pipelines: observation-level interfaces, constraint/prior/weight-level interfaces, solver-level interfaces, representation-level interfaces, and system-level integration interfaces. Based on this taxonomy, representative approaches are comparatively analyzed for accuracy, robustness, efficiency, and deployability. In addition, this review summarizes common design principles, including geometric consistency constraints, error propagation characteristics, and typical failure modes, and further discusses open challenges and future directions such as lightweight deployment, cross-domain adaptation, dynamic map modeling, and long-term consistency maintenance. This review aims to provide a structured reference for the analysis, design, and deployment of learning-enhanced VSLAM systems.

Keywords

Visual SLAM; deep learning; localization and mapping; autonomous navigation
  • 135

    View

  • 28

    Download

  • 0

    Like

Share Link