Skip to main content
QUICK REVIEW

[Paper Review] Comparing Human and Machine Deepfake Detection with Affective and Holistic Processing.

Matthew Groh, Ziv Epstein|arXiv (Cornell University)|May 13, 2021
Generative Adversarial Networks and Image Synthesis72 references4 citations
TL;DR

This study compares human and AI-based deepfake detection using 15,016 participants across three online experiments. It finds that humans and state-of-the-art models perform similarly in accuracy but make different errors; combining human judgment with model predictions improves detection when the model is correct, but model errors harm human performance. Emotional states and holistic facial processing significantly affect human detection, suggesting psychological and perceptual factors are key defenses against deepfakes.

ABSTRACT

The recent emergence of deepfake videos leads to an important societal question: how can we know if a video that we watch is real or fake? In three online studies with 15,016 participants, we present authentic videos and deepfakes and ask participants to identify which is which. We compare the performance of ordinary participants against the leading computer vision deepfake detection model and find them similarly accurate while making different kinds of mistakes. Together, participants with access to the model's prediction are more accurate than either alone, but inaccurate model predictions often decrease participants' accuracy. We embed randomized experiments and find: incidental anger decreases participants' performance and obstructing holistic visual processing of faces also hinders participants' performance while mostly not affecting the model's. These results suggest that considering emotional influences and harnessing specialized, holistic visual processing of ordinary people could be promising defenses against machine-manipulated media.

Motivation & Objective

  • To investigate how ordinary people compare to state-of-the-art computer vision models in detecting deepfake videos.
  • To examine the influence of incidental emotions and holistic visual processing on human deepfake detection performance.
  • To assess whether combining human judgment with AI model predictions enhances overall detection accuracy.
  • To identify psychological and perceptual factors that could serve as defenses against machine-manipulated media.

Proposed method

  • Conducted three online studies with 15,016 participants, showing authentic videos and deepfakes and asking them to identify which is real.
  • Compared human detection accuracy against predictions from a leading computer vision deepfake detection model.
  • Embedded randomized experiments to manipulate incidental anger and obstruct holistic visual processing of faces.
  • Measured performance changes under emotional and perceptual interference conditions.
  • Analyzed interaction effects between human decisions and model predictions, especially when model predictions were incorrect.
  • Used statistical modeling to assess the impact of emotional states and visual processing constraints on detection accuracy.

Experimental results

Research questions

  • RQ1How do ordinary humans compare in accuracy to state-of-the-art deepfake detection models in identifying manipulated videos?
  • RQ2How does incidental anger affect human performance in detecting deepfakes compared to model performance?
  • RQ3What is the impact of obstructing holistic facial processing on human deepfake detection accuracy, and does this affect the model similarly?
  • RQ4Can combining human judgment with AI model predictions improve overall detection accuracy, and under what conditions?
  • RQ5To what extent do emotional and perceptual factors moderate human susceptibility to deepfake deception?

Key findings

  • Humans and the leading computer vision deepfake detection model achieved similarly high levels of accuracy in identifying deepfakes.
  • Incidental anger significantly decreased human detection performance, while having little to no effect on the model’s accuracy.
  • Obstructing holistic visual processing of faces impaired human detection performance but did not significantly affect the model’s performance.
  • When the model provided correct predictions, combining its output with human judgment increased overall detection accuracy beyond either alone.
  • Inaccurate model predictions often reduced human detection accuracy, indicating that overreliance on AI can be detrimental.
  • The results suggest that affective states and holistic visual processing are critical, underappreciated factors in human deepfake detection.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.