Skip to main content
QUICK REVIEW

[论文解读] Improving Perceptual Quality by Phone-Fortified Perceptual Loss for Speech Enhancement.

Tsun-An Hsieh, Cheng Yu|arXiv (Cornell University)|Oct 28, 2020
Speech and Audio Processing参考文献 28被引用 14
一句话总结

本文提出了一种语音强化感知损失(PFP损失),通过将wav2vec表征中的语音信息整合到损失函数中,以提升语音增强的感知质量。在Voice Bank-DEMAND数据集上,通过在训练中利用语音结构,PFP损失在标准化的质量和可懂度指标上均优于逐点损失或信号级损失。

ABSTRACT

Speech enhancement (SE) aims to improve speech quality and intelligibility, which are both related to a smooth transition in speech segments that may carry linguistic information, e.g. phones and syllables. In this study, we took phonetic characteristics into account in the SE training process. Hence, we designed a phone-fortified perceptual (PFP) loss, and the training of our SE model was guided by PFP loss. In PFP loss, phonetic characteristics are extracted by wav2vec, an unsupervised learning model based on the contrastive predictive coding (CPC) criterion. Different from previous deep-feature-based approaches, the proposed approach explicitly uses the phonetic information in the deep feature extraction process to guide the SE model training. To test the proposed approach, we first confirmed that the wav2vec representations carried clear phonetic information using a t-distributed stochastic neighbor embedding (t-SNE) analysis. Next, we observed that the proposed PFP loss was more strongly correlated with the perceptual evaluation metrics than point-wise and signal-level losses, thus achieving higher scores for standardized quality and intelligibility evaluation metrics in the Voice Bank-DEMAND dataset.

研究动机与目标

  • 通过将语音结构整合到损失函数中,提升语音增强性能。
  • 探究深层特征中的语音信息是否能增强语音增强中的感知质量。
  • 开发一种更符合人类感知评估指标的损失函数。
  • 验证wav2vec表征是否编码了与语音增强相关的有意义的语音内容。
  • 证明PFP损失在感知质量和可懂度方面优于传统损失函数。

提出的方法

  • PFP损失被设计为利用基于对比预测编码(CPC)的无监督模型wav2vec提取的语音特征。
  • 通过wav2vec从干净语音中提取语音表征,并用于指导语音增强模型的训练。
  • 该方法使用t-SNE可视化确认wav2vec特征携带清晰的语音结构。
  • PFP损失采用端到端方式训练,模型在最小化感知损失的同时保留语音内容。
  • 该方法显式地将音素级语言信息整合到损失函数中,与以往基于深层特征的方法不同。
  • 在Voice Bank-DEMAND数据集上,使用标准化的感知指标对模型进行评估。

实验结果

研究问题

  • RQ1通过wav2vec提取的语音信息能否提升语音增强性能?
  • RQ2PFP损失是否与人类感知评估指标的相关性强于逐点损失或信号级损失?
  • RQ3在损失函数中整合语音结构是否能带来更高的感知质量和可懂度得分?
  • RQ4wav2vec表征是否包含适合语音增强的判别性语音信息?
  • RQ5PFP损失在感知质量提升方面与传统损失函数相比表现如何?

主要发现

  • t-SNE可视化证实wav2vec表征携带清晰的语音信息,支持其在PFP损失中的应用。
  • PFP损失与感知评估指标的相关性强于逐点损失或信号级损失。
  • 在Voice Bank-DEMAND数据集上,PFP损失在标准化的质量和可懂度评估指标上优于基线损失。
  • 在损失函数中整合语音结构可提升增强语音的感知质量。
  • 客观指标测量表明,所提方法在感知质量方面优于传统损失函数。
  • 结果表明,通过PFP损失实现的语音感知训练可使语音增强性能超越仅基于信号级优化的方法。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。