[Paper Review] Morphing Attack Detection -- Database, Evaluation Platform and Benchmarking
This paper introduces a new sequestered morphing attack detection (MAD) database and online evaluation platform to address critical gaps in benchmarking robustness and generalization of MAD algorithms. The dataset comprises 150 diverse subjects with morphed images created via careful subject selection and post-processing, including print-and-scan simulation to remove digital artifacts. Key results show that even state-of-the-art methods achieve BPCER100 above 90% on printed-and-scanned images, highlighting significant performance drops and the urgent need for more robust detection algorithms in real-world operational scenarios.
Morphing attacks have posed a severe threat to Face Recognition System (FRS). Despite the number of advancements reported in recent works, we note serious open issues such as independent benchmarking, generalizability challenges and considerations to age, gender, ethnicity that are inadequately addressed. Morphing Attack Detection (MAD) algorithms often are prone to generalization challenges as they are database dependent. The existing databases, mostly of semi-public nature, lack in diversity in terms of ethnicity, various morphing process and post-processing pipelines. Further, they do not reflect a realistic operational scenario for Automated Border Control (ABC) and do not provide a basis to test MAD on unseen data, in order to benchmark the robustness of algorithms. In this work, we present a new sequestered dataset for facilitating the advancements of MAD where the algorithms can be tested on unseen data in an effort to better generalize. The newly constructed dataset consists of facial images from 150 subjects from various ethnicities, age-groups and both genders. In order to challenge the existing MAD algorithms, the morphed images are with careful subject pre-selection created from the contributing images, and further post-processed to remove morphing artifacts. The images are also printed and scanned to remove all digital cues and to simulate a realistic challenge for MAD algorithms. Further, we present a new online evaluation platform to test algorithms on sequestered data. With the platform we can benchmark the morph detection performance and study the generalization ability. This work also presents a detailed analysis on various subsets of sequestered data and outlines open challenges for future directions in MAD research.
Motivation & Objective
- Address the lack of standardized, diverse, and realistic benchmarks for Morphing Attack Detection (MAD) in face recognition systems.
- Overcome generalization challenges in MAD algorithms by providing a sequestered dataset for testing on unseen data.
- Simulate real-world operational conditions in Automated Border Control (ABC) by including print-and-scan image processing pipelines.
- Enable continuous benchmarking of MAD algorithms through a dedicated online evaluation platform.
- Investigate the impact of demographic factors (age, gender, ethnicity) and morphing parameters on MAD performance.
Proposed method
- Constructed a new sequestered dataset of 150 subjects with diverse ethnicity, age, and gender to improve demographic generalization.
- Generated morphed images using advanced morphing techniques (e.g., GIMP GAP, Sqirlz Morph, FaceFusion) with careful subject pre-selection to ensure high-quality composites.
- Applied post-processing to remove visible morphing artifacts and simulate real-world conditions by printing and scanning images to eliminate digital cues.
- Developed an online evaluation platform to enable blind testing of MAD algorithms on sequestered data, supporting continuous benchmarking.
- Evaluated algorithms using standard metrics: Equal Error Rate (EER), BPCER at 10%, 20%, and 100% false acceptance rates, and rejection rates.
- Conducted ablation studies across subsets with varying morphing factors (0.3 and 0.5) to analyze algorithm robustness under different attack intensities.
Experimental results
Research questions
- RQ1How does the performance of existing MAD algorithms degrade when tested on unseen, sequestered data under realistic print-and-scan conditions?
- RQ2To what extent do demographic factors (age, gender, ethnicity) and morphing intensity influence the detectability of morphed images?
- RQ3Can the proposed online evaluation platform effectively support continuous benchmarking and generalization testing of MAD algorithms?
- RQ4What is the performance gap between digital and re-digitized (print-and-scan) morphed images across different MAD approaches?
- RQ5How do hybrid or multi-modal detection strategies compare to single-method approaches in detecting high-quality morphed images?
Key findings
- The best-performing D-MAD method (DFR) achieved a BPCER100 of 19.70% on print-and-scan images, indicating a 19.7% false acceptance rate at 100% false acceptance threshold, which is far from operational requirements.
- For S-MAD algorithms, BPCER100 exceeded 90% across all tested methods (e.g., 95.58% for WL, 100% for Deep-S-MAD), demonstrating severe performance degradation under print-and-scan conditions.
- A consistent performance drop was observed in all MAD algorithms when transitioning from digital to print-and-scan images, with BPCER100 increasing by up to 10 percentage points for some methods.
- The morphing factor (0.3 vs. 0.5) had a measurable but limited impact on detection performance, with slightly better results at lower morphing intensity.
- No significant improvement was observed across different morphing methods or detection techniques, indicating that current approaches lack robustness to realistic image degradation.
- The results confirm that cross-database training and testing are essential for improving generalization, as algorithms trained on one dataset fail on unseen, sequestered data.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.