[Paper Review] RGB-D And Thermal Sensor Fusion: A Systematic Literature Review
This systematic literature review synthesizes state-of-the-art techniques for fusing RGB-D and thermal sensor data, evaluating methods across calibration, 3D reconstruction, segmentation, and deep learning-based fusion. It identifies a critical gap in publicly available tri-modal datasets and highlights that while deep learning improves fusion performance, real-time processing remains challenging due to high computational demands, especially in middle fusion architectures.
In the last decade, the computer vision field has seen significant progress in multimodal data fusion and learning, where multiple sensors, including depth, infrared, and visual, are used to capture the environment across diverse spectral ranges. Despite these advancements, there has been no systematic and comprehensive evaluation of fusing RGB-D and thermal modalities to date. While autonomous driving using LiDAR, radar, RGB, and other sensors has garnered substantial research interest, along with the fusion of RGB and depth modalities, the integration of thermal cameras and, specifically, the fusion of RGB-D and thermal data, has received comparatively less attention. This might be partly due to the limited number of publicly available datasets for such applications. This paper provides a comprehensive review of both, state-of-the-art and traditional methods used in fusing RGB-D and thermal camera data for various applications, such as site inspection, human tracking, fault detection, and others. The reviewed literature has been categorised into technical areas, such as 3D reconstruction, segmentation, object detection, available datasets, and other related topics. Following a brief introduction and an overview of the methodology, the study delves into calibration and registration techniques, then examines thermal visualisation and 3D reconstruction, before discussing the application of classic feature-based techniques as well as modern deep learning approaches. The paper concludes with a discourse on current limitations and potential future research directions. It is hoped that this survey will serve as a valuable reference for researchers looking to familiarise themselves with the latest advancements and contribute to the RGB-DT research field.
Motivation & Objective
- To provide a comprehensive, systematic review of RGB-D and thermal sensor fusion techniques across multiple applications.
- To identify and analyze key challenges in sensor calibration, data alignment, and fusion strategies for RGB-DT modalities.
- To evaluate the performance of traditional feature-based and modern deep learning-based fusion approaches.
- To highlight the scarcity of publicly available tri-modal (RGB-D-T) datasets as a major bottleneck in advancing the field.
- To outline future research directions, particularly in real-time processing and robustness of fusion models.
Proposed method
- A systematic literature review was conducted using the PRISMA framework to identify peer-reviewed studies on RGB-D and thermal sensor fusion.
- The review categorized studies by technical domains: 3D reconstruction, segmentation, object detection, calibration, and dataset availability.
- Techniques were analyzed across three fusion levels: early, late, and middle fusion, with emphasis on deep learning architectures.
- The study evaluated both classical feature-based methods and modern deep learning models, including CNNs and Visual Transformers.
- Preprocessing steps such as image alignment, noise reduction, and thermal visualization were assessed for their impact on feature visibility.
- The analysis focused on computational efficiency, real-time performance, and generalization capabilities of fusion models.
Experimental results
Research questions
- RQ1What are the dominant approaches for fusing RGB-D and thermal sensor data in computer vision applications?
- RQ2How do calibration and registration techniques affect the accuracy of multi-modal perception systems?
- RQ3What are the performance trade-offs between early, late, and middle fusion strategies in RGB-DT systems?
- RQ4Why is the lack of publicly available tri-modal datasets a limiting factor in advancing RGB-DT research?
- RQ5What are the key challenges in achieving real-time, accurate, and robust sensor fusion using deep learning?
Key findings
- Deep learning-based fusion methods outperform traditional feature-based approaches in accuracy and robustness for object detection and segmentation tasks.
- Middle fusion strategies, while more accurate, often result in frame rates below 5 FPS, making real-time deployment difficult.
- Despite advances, no existing model has yet achieved effective tri-modal fusion of RGB, depth, and thermal data using Visual Transformers or similar advanced architectures.
- The scarcity of publicly available datasets—only VDT-2048 is identified as a suitable tri-modal dataset—significantly hinders progress in the field.
- Preprocessing steps such as image alignment and thermal visualization are critical for enhancing feature visibility in deep neural networks.
- Current Presentation Attack Detection (PAD) methods show bias toward training data, and generalization remains a key challenge, especially for unknown attack types.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.