[Paper Review] A Large-Scale, Time-Synchronized Visible and Thermal Face Dataset
This paper introduces the ARL-VTF dataset, the largest publicly available time-synchronized visible and long-wave infrared (LWIR) thermal face dataset, comprising over 500,000 images from 395 subjects. It enables benchmarking of thermal face landmark detection and thermal-to-visible face verification, achieving state-of-the-art AUC scores of 99.28% with SAGAN-based synthesis and 99.63% with ground-truth visible probes.
Thermal face imagery, which captures the naturally emitted heat from the face, is limited in availability compared to face imagery in the visible spectrum. To help address this scarcity of thermal face imagery for research and algorithm development, we present the DEVCOM Army Research Laboratory Visible-Thermal Face Dataset (ARL-VTF). With over 500,000 images from 395 subjects, the ARL-VTF dataset represents, to the best of our knowledge, the largest collection of paired visible and thermal face images to date. The data was captured using a modern long wave infrared (LWIR) camera mounted alongside a stereo setup of three visible spectrum cameras. Variability in expressions, pose, and eyewear has been systematically recorded. The dataset has been curated with extensive annotations, metadata, and standardized protocols for evaluation. Furthermore, this paper presents extensive benchmark results and analysis on thermal face landmark detection and thermal-to-visible face verification by evaluating state-of-the-art models on the ARL-VTF dataset.
Motivation & Objective
- To address the scarcity of high-quality, time-synchronized visible and thermal face imagery for research.
- To support development of robust face recognition models under unconstrained conditions such as pose variation, expression, and occlusion.
- To provide standardized protocols for training and evaluating thermal-to-visible face verification and landmark detection models.
- To benchmark state-of-the-art deep learning models on thermal face analysis tasks using a large-scale, curated dataset.
Proposed method
- Simultaneous capture of visible and LWIR thermal images using a stereo visible camera setup and a long-wave infrared (LWIR) camera.
- Systematic recording of three image sequences per subject: baseline, expression, and off-pose, with a fourth for eyewear conditions.
- Annotation of head pose, eyewear, face bounding boxes, and 6 facial landmarks for all images.
- Application of deep learning models including VGG-Face, Pix2Pix, GANVFS, and SAGAN for cross-modal face verification and image synthesis.
- Use of conditional GANs (e.g., SAGAN) to synthesize visible-like images from thermal inputs to improve cross-spectrum matching.
- Evaluation using standardized protocols with metrics including AUC, EER, and landmark detection accuracy across varying conditions.
Experimental results
Research questions
- RQ1How does thermal-to-visible face verification performance vary across different pose and expression conditions?
- RQ2To what extent does eyewear occlusion degrade thermal-to-visible face verification accuracy, especially in the thermal domain?
- RQ3How effective are GAN-based image synthesis methods in improving cross-spectral face verification performance?
- RQ4What is the impact of pose variation on thermal face landmark detection and verification accuracy?
- RQ5How do different synthesis methods (e.g., Pix2Pix, SAGAN) compare in generating realistic visible-like images from thermal inputs?
Key findings
- The SAGAN-based synthesis method achieved an AUC of 99.28% on thermal-to-visible face verification, significantly outperforming the raw baseline (61.37% AUC).
- The ground-truth visible probe baseline achieved 99.06% AUC on baseline sequences, but performance dropped to 75.76% on off-pose sequences, highlighting pose as a major challenge.
- Performance degradation due to expression was observed, with SAGAN’s AUC dropping from 99.28% to 98.46% on expressive faces.
- The Equal Error Rate (EER) for the SAGAN model was 3.97%, indicating high robustness, while the Pix2Pix model had an EER of 33.8% on pose variations.
- Thermal images with eyeglasses caused significant performance drops due to heat absorption in lenses, which occludes thermal signatures and degrades matching accuracy.
- The dataset’s standardized protocols and extensive annotations enable reliable benchmarking of thermal face analysis models across diverse real-world conditions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.