[论文解读] ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech
本文介绍了 ASVspoof 2019 数据库,这是一个大规模公开数据集,涵盖在逻辑访问和物理访问场景下生成的合成语音、转换语音和重放语音。该研究通过自动化 ASV 系统和反制系统以及人工评估,对最先进的欺骗攻击进行了评估,结果表明,某些合成语音甚至对人类听者也难以与真实语音区分,凸显了语音验证系统中的关键漏洞。
Automatic speaker verification (ASV) is one of the most natural and convenient means of biometric person recognition. Unfortunately, just like all other biometric systems, ASV is vulnerable to spoofing, also referred to as "presentation attacks." These vulnerabilities are generally unacceptable and call for spoofing countermeasures or "presentation attack detection" systems. In addition to impersonation, ASV systems are vulnerable to replay, speech synthesis, and voice conversion attacks. The ASVspoof 2019 edition is the first to consider all three spoofing attack types within a single challenge. While they originate from the same source database and same underlying protocol, they are explored in two specific use case scenarios. Spoofing attacks within a logical access (LA) scenario are generated with the latest speech synthesis and voice conversion technologies, including state-of-the-art neural acoustic and waveform model techniques. Replay spoofing attacks within a physical access (PA) scenario are generated through carefully controlled simulations that support much more revealing analysis than possible previously. Also new to the 2019 edition is the use of the tandem detection cost function metric, which reflects the impact of spoofing and countermeasures on the reliability of a fixed ASV system. This paper describes the database design, protocol, spoofing attack implementations, and baseline ASV and countermeasure results. It also describes a human assessment on spoofed data in logical access. It was demonstrated that the spoofing data in the ASVspoof 2019 database have varied degrees of perceived quality and similarity to the target speakers, including spoofed data that cannot be differentiated from bona-fide utterances even by human subjects.
研究动机与目标
- 开发一个全面的基准,用于评估自动语音验证(ASV)系统中的欺骗反制措施。
- 解决 ASV 系统对三种主要欺骗攻击的脆弱性:文本到语音合成、语音转换和重放攻击。
- 提供一个标准化的、公开可用的数据集,包含真实标签和元数据,以支持学术界和工业界的研究。
- 通过自动化指标(EER)和人工评估,评估 ASV 和反制系统的表现。
- 在受控声学条件下,实现对欺骗攻击有效性的系统性分析,特别是在物理访问场景中。
提出的方法
- 该数据库包含使用最先进的神经声学和波形模型生成的欺骗语音,包括用于文本到语音合成和语音转换的 Tacotron2、WaveNet 和 WaveRNN。
- 重放攻击在受控声学环境中模拟,通过精确的麦克风和扬声器位置布置,以实现对物理访问条件的详细分析。
- 定义了两种使用场景:用于合成/转换语音的逻辑访问(LA)和用于重放攻击的物理访问(PA),每种场景均有独立的协议和评估指标。
- 引入了并联检测代价函数(t-DCF)作为新指标,用于评估欺骗和反制措施对 ASV 可靠性的影响。
- 在数据集上训练并评估了基线 ASV 和反制(CM)系统,使用等错误率(EER)和 t-DCF 来评估性能。
- 在 LA 子集上开展了人工评估,以评估欺骗语音的感知质量和与目标说话人的相似度。
实验结果
研究问题
- RQ1最先进的文本到语音和语音转换系统在生成与真实语音难以区分的欺骗语音方面效果如何?
- RQ2在受控声学条件下,物理访问场景中的重放攻击在多大程度上损害了 ASV 系统的可靠性?
- RQ3自动化 ASV 和反制系统在新数据集上的表现如何?其结果与人类感知的相关性如何?
- RQ4波形生成技术(例如 Griffin-Lim 与 WaveRNN)对欺骗语音的感知质量和可检测性有何影响?
- RQ5新引入的 t-DCF 指标在多大程度上反映了 ASV 系统在欺骗攻击下的实际可靠性?
主要发现
- 使用 WaveRNN 进行波形生成的系统 A10 所生成的欺骗语音,即使对人类听者也难以与真实语音区分,感知质量和相似度无显著差异(p=0.012 和 p=0.81,分别)。
- 使用 Griffin-Lim 进行波形生成的系统 A11 生成的合成语音质量较低,且人类听者能显著更易检测到,表明波形生成技术对欺骗质量具有决定性影响。
- A10 与 A11 在可检测性上的差异具有统计显著性(p<0.001),证实波形生成技术同时影响感知质量和可检测性。
- 人类评估者比自动化系统更容易检测到 A13 和 A17 生成的欺骗语音,表明人类感知与基线 ASV/CM 性能之间存在脱节。
- t-DCF 指标有效捕捉了欺骗攻击对 ASV 系统可靠性的影响,提供了比 EER 单独评估更真实的评估结果。
- ASVspoof 2019 数据库包含在感知上与真实语音无法区分的欺骗语音,证明了现代 TTS 和 VC 系统在绕过人类和机器检测方面的先进能力。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。