[Paper Review] Testing Human Ability To Detect Deepfake Images of Human Faces
This study evaluates human ability to detect StyleGAN2-generated deepfake images of human faces using an online survey with 280 participants across four groups, including control and intervention conditions. Despite high confidence, participants achieved only 62% accuracy—slightly above chance—with no intervention significantly improving detection, highlighting a critical vulnerability in human perception of synthetic media.
Deepfakes are computationally-created entities that falsely represent reality. They can take image, video, and audio modalities, and pose a threat to many areas of systems and societies, comprising a topic of interest to various aspects of cybersecurity and cybersafety. In 2020 a workshop consulting AI experts from academia, policing, government, the private sector, and state security agencies ranked deepfakes as the most serious AI threat. These experts noted that since fake material can propagate through many uncontrolled routes, changes in citizen behaviour may be the only effective defence. This study aims to assess human ability to identify image deepfakes of human faces (StyleGAN2:FFHQ) from nondeepfake images (FFHQ), and to assess the effectiveness of simple interventions intended to improve detection accuracy. Using an online survey, 280 participants were randomly allocated to one of four groups: a control group, and 3 assistance interventions. Each participant was shown a sequence of 20 images randomly selected from a pool of 50 deepfake and 50 real images of human faces. Participants were asked if each image was AI-generated or not, to report their confidence, and to describe the reasoning behind each response. Overall detection accuracy was only just above chance and none of the interventions significantly improved this. Participants' confidence in their answers was high and unrelated to accuracy. Assessing the results on a per-image basis reveals participants consistently found certain images harder to label correctly, but reported similarly high confidence regardless of the image. Thus, although participant accuracy was 62% overall, this accuracy across images ranged quite evenly between 85% and 30%, with an accuracy of below 50% for one in every five images. We interpret the findings as suggesting that there is a need for an urgent call to action to address this threat.
Motivation & Objective
- To assess human ability to detect deepfake images of human faces created with StyleGAN2.
- To evaluate the effectiveness of simple interventions in improving human detection accuracy.
- To examine the relationship between confidence levels and actual detection accuracy.
- To identify specific image characteristics that make deepfakes particularly difficult to detect.
Proposed method
- Participants were recruited via an online survey and randomly assigned to one of four groups: control or three intervention conditions.
- Each participant viewed 20 randomly selected images from a balanced set of 50 real and 50 deepfake images from the FFHQ dataset.
- Participants classified each image as AI-generated or real, reporting both their confidence and reasoning for each judgment.
- Detection accuracy was calculated per participant and per image, with confidence levels correlated to accuracy.
- Statistical analysis compared detection performance across groups to assess intervention effectiveness.
- Per-image analysis identified consistently misclassified images, revealing patterns in human difficulty with specific deepfake samples.
Experimental results
Research questions
- RQ1What is the baseline human detection accuracy for deepfake images of human faces using StyleGAN2?
- RQ2Do simple interventions significantly improve human ability to detect deepfake images of human faces?
- RQ3How does self-reported confidence correlate with actual detection accuracy in identifying deepfakes?
- RQ4Which specific deepfake images are most frequently misclassified by humans, and what features make them deceptive?
Key findings
- Overall detection accuracy was 62%, only slightly above chance performance, indicating poor human ability to reliably detect deepfake images.
- No significant improvement in detection accuracy was observed across any of the three intervention conditions compared to the control group.
- Participants reported high confidence in their judgments, but this confidence showed no meaningful correlation with actual accuracy.
- Detection accuracy varied widely per image, ranging from 30% to 85%, with one in five images achieving below 50% accuracy.
- Certain deepfake images were consistently misclassified across participants, suggesting specific visual artifacts or patterns that evade human perception.
- The findings indicate a systemic vulnerability in human perception, underscoring the urgent need for technical and educational interventions to counter deepfake threats.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.