[Paper Review] CASIA-SURF: A Dataset and Benchmark for Large-scale Multi-modal Face Anti-spoofing.
This paper introduces CASIA-SURF, the largest publicly available face anti-spoofing dataset with 1,000 subjects, 21,000 videos, and three modalities (RGB, Depth, IR), along with a standardized benchmark, evaluation protocol, and training/validation/testing splits. It proposes a multi-modal fusion method that reweights informative features while suppressing less useful ones, achieving strong performance on the new benchmark.
Face anti-spoofing is essential to prevent face recognition systems from a security breach. Much of the progresses have been made by the availability of face anti-spoofing benchmark datasets in recent years. However, existing face anti-spoofing benchmarks have limited number of subjects ($\le egmedspace170$) and modalities ($\leq egmedspace2$), which hinder the further development of the academic community. To facilitate future face anti-spoofing research, we introduce a large-scale multi-modal dataset, namely CASIA-SURF, which is the largest publicly available dataset for face anti-spoofing both in terms of subjects and visual modalities. Specifically, it consists of $1,000$ subjects with $21,000$ videos and each sample has $3$ modalities (i.e., RGB, Depth and IR). Associated with this dataset, we also provide concrete measurement set, evaluation protocol and training/validation/testing subsets, developing a new benchmark for face anti-spoofing. Moreover, we present a new multi-modal fusion method as a strong baseline, which performs feature re-weighting to select the more informative channel features while suppressing less useful ones for each modal. Extensive experiments have been conducted on the proposed dataset to verify its significance and generalization capability. Dataset is available at https://sites.google.com/qq.com/chalearnfacespoofingattackdete
Motivation & Objective
- To address the limited scale and modality diversity in existing face anti-spoofing benchmarks, which hinder progress in the field.
- To provide a large-scale, multi-modal dataset with 1,000 subjects and 21,000 videos across RGB, Depth, and IR modalities.
- To establish a standardized benchmark with defined training, validation, and testing subsets for consistent evaluation.
- To develop a strong multi-modal fusion baseline method that adaptively reweights informative features across modalities.
- To enable generalization and scalability studies in large-scale, multi-modal face anti-spoofing.
Proposed method
- The dataset, CASIA-SURF, is constructed with 1,000 subjects, each contributing multiple videos under controlled conditions across three visual modalities: RGB, Depth, and Infrared (IR).
- A standardized evaluation protocol is defined, including fixed training, validation, and testing splits to ensure reproducibility and fair comparison.
- A novel multi-modal fusion method is proposed that performs feature re-weighting per channel to emphasize informative features and suppress less useful ones across modalities.
- The method uses learnable attention mechanisms or similar re-weighting strategies to dynamically adjust feature importance during fusion.
- The benchmark supports training and evaluation of deep learning models on large-scale, multi-modal anti-spoofing tasks.
- The dataset is publicly released at https://sites.google.com/qq.com/chalearnfacespoofingattackdete to promote community-wide research.
Experimental results
Research questions
- RQ1How does model performance scale with increased subject and video diversity in face anti-spoofing datasets?
- RQ2What is the impact of incorporating multiple visual modalities (RGB, Depth, IR) on spoof detection accuracy?
- RQ3Can a learnable feature re-weighting mechanism improve multi-modal fusion performance compared to simple concatenation or averaging?
- RQ4How generalizable are models trained on CASIA-SURF to unseen subjects and spoofing attacks?
- RQ5What is the performance of a strong multi-modal fusion baseline on the new benchmark under standardized evaluation protocols?
Key findings
- CASIA-SURF is the largest publicly available face anti-spoofing dataset, with 1,000 subjects and 21,000 videos across three modalities (RGB, Depth, IR).
- The proposed multi-modal fusion method with feature re-weighting achieves superior performance by selectively emphasizing informative features and suppressing less useful ones.
- The benchmark’s standardized training, validation, and testing splits enable consistent and reproducible evaluation across different models.
- Extensive experiments confirm the dataset’s generalization capability and suitability for large-scale, multi-modal anti-spoofing research.
- The dataset and benchmark are publicly released to accelerate progress in face anti-spoofing research.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.