[Paper Review] Celeb-DF: A Large-scale Challenging Dataset for DeepFake Forensics
Celeb-DF introduces a large-scale, high-quality DeepFake video dataset (5,639 DeepFakes, over 2 million frames) to better evaluate detection methods, showing current detectors struggle with this higher-quality data.
AI-synthesized face-swapping videos, commonly known as DeepFakes, is an emerging problem threatening the trustworthiness of online information. The need to develop and evaluate DeepFake detection algorithms calls for large-scale datasets. However, current DeepFake datasets suffer from low visual quality and do not resemble DeepFake videos circulated on the Internet. We present a new large-scale challenging DeepFake video dataset, Celeb-DF, which contains 5,639 high-quality DeepFake videos of celebrities generated using improved synthesis process. We conduct a comprehensive evaluation of DeepFake detection methods and datasets to demonstrate the escalated level of challenges posed by Celeb-DF.
Motivation & Objective
- Motivate the need for a larger, higher-quality DeepFake video dataset that better matches realistic internet content.
- Create Celeb-DF with improved synthesis to reduce artifacts seen in prior datasets.
- Provide a comprehensive evaluation of current DeepFake detection methods on Celeb-DF and existing datasets to assess real-world challenges.
Proposed method
- Develop an improved DeepFake synthesis pipeline producing 256x256 face regions, higher visual quality, and fewer artifacts.
- Apply color augmentation and color transfer to reduce color mismatch between donor and target faces.
- Improve face masks to cover complete facial regions with smooth boundaries.
- Incorporate Kalman smoothing to temporal landmarks to reduce frame-to-frame flickering.
- Quantitatively assess visual quality using Mask-SSIM focusing on head regions.
- Evaluate detection methods using frame-level AUC on multiple datasets including Celeb-DF.
Experimental results
Research questions
- RQ1How does the Celeb-DF dataset compare in visual quality to previous DeepFake datasets?
- RQ2How do current DeepFake detection methods perform on Celeb-DF relative to earlier datasets?
- RQ3What is the impact of video compression on detection performance across state-of-the-art detectors?
Key findings
- Celeb-DF contains 5,639 DeepFake videos (over 2,000,000 frames) and 590 real videos from 59 celebrities.
- The Celeb-DF synthesis pipeline yields higher visual quality with fewer artifacts, as evidenced by higher Mask-SSIM scores ( Celeb-DF: 0.92 -SSIM listed in Table 2).
- Across evaluated detectors, Celeb-DF is generally the most challenging dataset, with lower average frame-level AUC compared to older datasets.
- A recent method (DSP-FWA) achieves the top performance among tested detectors, with about 87.4% in the study’s summary (overall performance across datasets).
- Compression experiments show detector performance degrades with higher H.264 compression, though some models (e.g., Xception variants) remain relatively robust.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.