[论文解读] Can we steal your vocal identity from the Internet?: Initial investigation of cloning Obama's voice using GAN, WaveNet and low-quality found data
本文研究了利用互联网上公开获取的低质量音频,通过生成对抗网络(GAN)语音增强模型提升音质后,克隆巴拉克·奥巴马声音的可行性。实验采用先进的文本到语音(TTS)与语音转换(VC)系统进行训练。结果表明,语音质量与说话人相似度显著提升,但感知伪影及较低的说话人相似度评分表明,当前系统仍无法达到类人自然度与身份保真度。
Thanks to the growing availability of spoofing databases and rapid advances in using them, systems for detecting voice spoofing attacks are becoming more and more capable, and error rates close to zero are being reached for the ASVspoof2015 database. However, speech synthesis and voice conversion paradigms that are not considered in the ASVspoof2015 database are appearing. Such examples include direct waveform modelling and generative adversarial networks. We also need to investigate the feasibility of training spoofing systems using only low-quality found data. For that purpose, we developed a generative adversarial network-based speech enhancement system that improves the quality of speech data found in publicly available sources. Using the enhanced data, we trained state-of-the-art text-to-speech and voice conversion models and evaluated them in terms of perceptual speech quality and speaker similarity. The results show that the enhancement models significantly improved the SNR of low-quality degraded data found in publicly available sources and that they significantly improved the perceptual cleanliness of the source speech without significantly degrading the naturalness of the voice. However, the results also show limitations when generating speech with the low-quality found data.
研究动机与目标
- 评估是否可利用互联网上公开获取的低质量音频训练出有效的语音克隆系统。
- 开发一种鲁棒的基于GAN的语音增强模型,以改善来自非受控环境的劣质音频。
- 评估经增强的低质量数据生成的合成语音在感知质量与说话人相似度方面的表现。
- 测试这些合成语音在绕过现有反欺骗检测机制方面的有效性。
- 确定下一代ASVspoof数据集是否应包含来自公开渠道的低质量数据。
提出的方法
- 训练一种改进的SEGAN-based语音增强模型,以提升来自公开来源的低质量、噪声音频的信噪比(SNR)与感知质量。
- 生成器采用跳跃连接(skip connection)以学习残差增强,降低从零开始生成清晰语音的负担。
- 生成器在早期训练阶段使用基线增强模型进行预训练,其损失函数相对于基线输出而非真实干净语音计算。
- 使用增强后的音频训练文本到语音(TTS)与语音转换(VC)模型,涵盖单语与双语配置。
- 采用CQCC-GMM反欺骗分类器评估合成语音的可检测性,该分类器在ASVspoof2015与VCC2016数据集上进行训练。
- 通过MOS(平均意见得分)听音测试进行感知评估,将合成语音与复制合成语音及自然语音进行对比。
实验结果
研究问题
- RQ1基于GAN的语音增强系统能否有效提升来自非受控环境的低质量公开音频的质量?
- RQ2从增强后的低质量数据生成的合成语音在感知质量与目标说话人相似度方面能达到何种程度?
- RQ3在增强后的公开数据上训练的现代TTS与语音转换系统,对现有反欺骗检测机制的规避效果如何?
- RQ4数据质量与训练数据选择对语音克隆系统性能有何影响?
- RQ5当在大规模多样化数据集上训练时,语音增强模型是否能在不同录音条件下实现良好泛化?
主要发现
- 基于GAN的语音增强模型显著提升了低质量公开音频的信噪比(SNR),并增强了感知上的清晰度,同时未损害自然度。
- 感知评估显示,增强后的音频在质量与自然度方面的MOS得分高于未经处理的低质量数据。
- 在增强数据上训练的TTS与VC系统,其感知质量优于在原始低质量数据上训练的系统,但说话人相似度仍较低(例如,复制合成语音为2.63 MOS,WaveNet-based TTS为2.45 MOS)。
- 添加混响可提升WaveNet生成语音的感知质量,但对说话人相似度的提升不显著。
- 基于CQCC-GMM的反欺骗检测机制在VCC2016数据集上达到高EER(如TT3在VCC2016集上EER为0.79%),表明即使经过先进增强与合成,合成语音仍可被有效检测。
- VC系统在质量与相似度方面均优于TTS系统,VC2(单语)在VCC2016训练的反检测器上EER为0.00%,表明合成语音具有强可检测性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。