Skip to main content
QUICK REVIEW

[Paper Review] Warning: Humans Cannot Reliably Detect Speech Deepfakes

Kimberly T. Mai, Sergi D. Bray|arXiv (Cornell University)|Jan 19, 2023
Speech Recognition and Synthesis4 citations
TL;DR

This study evaluates human ability to detect speech deepfakes using a controlled online experiment with 529 participants across English and Mandarin. Despite awareness and exposure to examples, listeners only detected deepfakes with 73% accuracy, showing no significant language difference, and familiarity only slightly improved performance—highlighting the urgent need for automated detection systems over human reliance.

ABSTRACT

Speech deepfakes are artificial voices generated by machine learning models. Previous literature has highlighted deepfakes as one of the biggest security threats arising from progress in artificial intelligence due to their potential for misuse. However, studies investigating human detection capabilities are limited. We presented genuine and deepfake audio to n = 529 individuals and asked them to identify the deepfakes. We ran our experiments in English and Mandarin to understand if language affects detection performance and decision-making rationale. We found that detection capability is unreliable. Listeners only correctly spotted the deepfakes 73% of the time, and there was no difference in detectability between the two languages. Increasing listener awareness by providing examples of speech deepfakes only improves results slightly. As speech synthesis algorithms improve and become more realistic, we can expect the detection task to become harder. The difficulty of detecting speech deepfakes confirms their potential for misuse and signals that defenses against this threat are needed.

Motivation & Objective

  • To assess human detection accuracy of speech deepfakes in English and Mandarin to evaluate language-dependent detection performance.
  • To investigate whether familiarizing listeners with deepfake examples improves detection capability.
  • To examine whether contextual cues or audio pairing (binary vs. unary) influence detection performance.
  • To understand the role of perceptual cues like unnaturalness in human deepfake detection.
  • To inform the development of robust defenses by identifying the limitations of human perception in detecting synthetic speech.

Proposed method

  • Conducted an online experiment with 529 participants across English and Mandarin-speaking groups.
  • Presented audio clips in two configurations: unary (single clip per trial) and binary (pair of clips, one real, one fake).
  • Used a deepfake voice synthesis model based on older, less advanced techniques to generate stimuli.
  • Provided randomized familiarization interventions with example deepfakes to assess impact on detection performance.
  • Collected confidence-adjusted accuracy scores to analyze decision reliability.
  • Analyzed detection performance across language, configuration, and familiarity conditions using statistical modeling.
Fig 1: Diagram of a typical generative speech synthesis model.
Fig 1: Diagram of a typical generative speech synthesis model.

Experimental results

Research questions

  • RQ1Can humans reliably distinguish genuine speech from speech deepfakes in English and Mandarin?
  • RQ2Does prior exposure to deepfake examples improve human detection accuracy?
  • RQ3Are there significant differences in detection performance between unary and binary audio presentation formats?
  • RQ4Do listeners rely on language-specific acoustic cues when detecting deepfakes?
  • RQ5How does the perceived naturalness of speech influence human detection decisions?

Key findings

  • Humans detected speech deepfakes with only 73% accuracy, indicating unreliable detection capability even under controlled conditions.
  • There was no significant difference in detection performance between English and Mandarin listeners, suggesting language does not affect detectability.
  • Familiarizing participants with deepfake examples improved detection accuracy only slightly, indicating limited effectiveness of awareness campaigns.
  • Listeners primarily relied on perceived unnaturalness as a cue, regardless of language, indicating a shared perceptual heuristic.
  • Performance did not improve with increased listening time or binary comparison, suggesting cognitive load and task design do not significantly aid detection.
  • The results imply that human-based detection is insufficient for real-world defense, necessitating automated detection systems.
Fig 2: Screenshots of the task interface.
Fig 2: Screenshots of the task interface.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.