Skip to main content
QUICK REVIEW

[论文解读] How Deep Are the Fakes? Focusing on Audio Deepfake: A Survey

Zahra Khanjani, Gabrielle Watson|arXiv (Cornell University)|Nov 28, 2021
Digital Media Forensic Detection参考文献 50被引用 6
一句话总结

本综述首次对音频深度伪造进行了全面的英文分析,聚焦于2016至2020年间的生成与检测方法。研究识别出生成对抗网络(GANs)、卷积神经网络(CNNs)和深度神经网络(DNNs)为核心技术,揭示了尽管威胁日益加剧,音频深度伪造检测仍存在关键研究空白,并呼吁紧急开发稳健的检测系统,以应对欺诈和虚假信息日益增长的滥用问题。

ABSTRACT

Deepfake is content or material that is synthetically generated or manipulated using artificial intelligence (AI) methods, to be passed off as real and can include audio, video, image, and text synthesis. This survey has been conducted with a different perspective compared to existing survey papers, that mostly focus on just video and image deepfakes. This survey not only evaluates generation and detection methods in the different deepfake categories, but mainly focuses on audio deepfakes that are overlooked in most of the existing surveys. This paper critically analyzes and provides a unique source of audio deepfake research, mostly ranging from 2016 to 2020. To the best of our knowledge, this is the first survey focusing on audio deepfakes in English. This survey provides readers with a summary of 1) different deepfake categories 2) how they could be created and detected 3) the most recent trends in this domain and shortcomings in detection methods 4) audio deepfakes, how they are created and detected in more detail which is the main focus of this paper. We found that Generative Adversarial Networks(GAN), Convolutional Neural Networks (CNN), and Deep Neural Networks (DNN) are common ways of creating and detecting deepfakes. In our evaluation of over 140 methods we found that the majority of the focus is on video deepfakes and in particular in the generation of video deepfakes. We found that for text deepfakes there are more generation methods but very few robust methods for detection, including fake news detection, which has become a controversial area of research because of the potential of heavy overlaps with human generation of fake content. This paper is an abbreviated version of the full survey and reveals a clear need to research audio deepfakes and particularly detection of audio deepfakes.

研究动机与目标

  • 为解决现有综述中对音频深度伪造关注不足的问题,这些综述主要聚焦于视频和图像深度伪造。
  • 系统分析音频深度伪造特有的生成与检测技术,填补关键研究空白。
  • 评估音频深度伪造检测领域的最新进展,突出其不足之处与研究需求。
  • 提供音频深度伪造方法的详细框架与分类体系,涵盖重放攻击、语音合成与语音转换。
  • 由于现实世界中滥用行为日益增加且检测能力不足,呼吁加大对音频深度伪造检测的研究投入。

提出的方法

  • 对2016至2020年间关于深度伪造的140余篇研究论文进行了系统性回顾,重点聚焦于音频特定方法。
  • 将音频深度伪造技术分类为三个子类别:重放攻击、语音合成与语音转换。
  • 使用等错误率(EER)、FID、PPL及感知质量评分等指标评估检测框架。
  • 分析用于生成与检测任务的深度学习架构,包括生成对抗网络(GANs)、卷积神经网络(CNNs)、深度神经网络(DNNs)和循环神经网络(RNNs)。
  • 提供对比表格(表1、2、3),总结关键框架、数据集、目标与性能指标。
  • 以ASVspoof2017作为基准数据集,评估重放攻击检测性能,利用深度卷积网络实现0%的等错误率(EER)。

实验结果

研究问题

  • RQ1音频深度伪造的主要技术类别与子类别是什么?它们在生成机制上如何不同?
  • RQ2检测音频深度伪造最有效的深度学习架构是什么?它们在性能上如何比较?
  • RQ3为何音频深度伪造检测相较于视频与图像深度伪造检测显著研究不足?
  • RQ4当前音频深度伪造检测方法在实际部署中存在哪些关键局限与不足?
  • RQ5近期框架如元学习生成对抗网络(meta-learning GANs)或神经渲染技术如何推动音频与视频深度伪造的生成与检测发展?

主要发现

  • 生成对抗网络(GANs)、卷积神经网络(CNNs)和深度神经网络(DNNs)是生成与检测音频深度伪造的主导架构。
  • 尽管威胁持续加剧,仅有极少数深度伪造研究聚焦于音频,且检测方法远落后于生成技术。
  • 深度卷积网络在ASVspoof2017数据集上对重放攻击检测实现了完美的0%等错误率(EER),优于以往方法。
  • 语音合成(TTS)与语音转换是音频深度伪造的主要子类别,近期框架已实现高保真度、唇音同步的音视频合成。
  • 本研究识别出在基于文本与基于音频的深度伪造检测方面,缺乏稳健且可泛化的检测方法,尤其是在区分AI生成内容与人为制造的虚假信息方面。
  • 综述表明,尽管视频深度伪造生成技术高度先进,但音频深度伪造检测仍严重滞后,造成严重的安全与社会风险。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。