Skip to main content
QUICK REVIEW

[论文解读] Contaminated speech training methods for robust DNN-HMM distant speech recognition

Mirco Ravanelli, Maurizio Omologo|arXiv (Cornell University)|Oct 10, 2017
Speech Recognition and Synthesis被引用 6
一句话总结

本文提出三种新颖方法——非对称上下文窗口、近讲监督和近讲预训练——以提升使用污染语音的远距离语音识别中DNN-HMM声学建模性能。这些技术在模拟和真实世界数据上均将词错误率降低超过15%,显著提升了混响和噪声环境下的鲁棒性。

ABSTRACT

Despite the significant progress made in the last years, state-of-the-art speech recognition technologies provide a satisfactory performance only in the close-talking condition. Robustness of distant speech recognition in adverse acoustic conditions, on the other hand, remains a crucial open issue for future applications of human-machine interaction. To this end, several advances in speech enhancement, acoustic scene analysis as well as acoustic modeling, have recently contributed to improve the state-of-the-art in the field. One of the most effective approaches to derive a robust acoustic modeling is based on using contaminated speech, which proved helpful in reducing the acoustic mismatch between training and testing conditions. In this paper, we revise this classical approach in the context of modern DNN-HMM systems, and propose the adoption of three methods, namely, asymmetric context windowing, close-talk based supervision, and close-talk based pre-training. The experimental results, obtained using both real and simulated data, show a significant advantage in using these three methods, overall providing a 15% error rate reduction compared to the baseline systems. The same trend in performance is confirmed either using a high-quality training set of small size, and a large one.

研究动机与目标

  • 解决在混响和背景噪声等不利声学条件下远距离语音识别的长期挑战。
  • 通过利用高质量近讲数据增强远距离语音训练中的监督和初始化,改进DNN-HMM声学建模。
  • 研究在现代DNN-HMM系统中使用污染语音训练的有效性,特别是优化上下文处理和预训练策略的效果。
  • 在模拟和真实世界远距离语音数据上验证所提方法,包括具有挑战性的真实混响和噪声条件。

提出的方法

  • 提出一种非对称上下文窗口(ACW),将时间焦点向未来帧偏移,减轻混响对DNN输入表征的负面影响。
  • 提出从近讲数据继承高精度强制对齐标签,用于监督远距离语音DNN训练,提升监督质量。
  • 用使用近讲数据的监督预训练替代标准无监督RBM预训练,实现远距离语音DNN的更优初始化。
  • 采用两阶段训练流程:首先在近讲数据上预训练DNN,然后在远距离污染数据上以较低的初始学习率微调。
  • 使用Kaldi自动语音识别工具包实现系统,并在电话环任务上评估性能,以隔离声学建模的影响。
  • 通过卷积测量的脉冲响应和在可控信噪比下添加噪声的方式实现语音污染,以模拟真实的远距离语音条件。

实验结果

研究问题

  • RQ1非对称上下文窗口是否能通过更好地处理混响影响,提升DNN-HMM在远距离语音识别中的性能?
  • RQ2从近讲数据继承高质量标签是否能加快远距离语音DNN训练的收敛速度并提升性能?
  • RQ3使用近讲数据的监督预训练是否能优于标准无监督RBM预训练,在远距离语音声学建模中表现更优?
  • RQ4所提方法在不同训练数据规模和真实声学条件下具有多大程度的泛化能力?

主要发现

  • 所提出的基于近讲的监督方法(CT-lab)在APASCI数据集上将词错误率降低10%,在Euronews数据集上降低12%,相比标准训练。
  • CT-lab方法加速了训练收敛,将训练轮次从15轮减少至12轮,从而将训练时间减少20%。
  • 基于近讲的监督预训练(CT-PT)实现了稳定提升,在APASCI数据集上分别将错误率降低1.6%(Sim-Rev)、1.2%(Real-Rev)和1.1%(Sim-Rev&Noise)。
  • ACW、CT-lab和CT-PT方法的组合在所有测试条件下相比基线系统实现了超过15%的相对错误率降低。
  • 性能提升在小规模(6小时)和大规模(100小时)训练集上均被观察到,证实了该方法的可扩展性和鲁棒性。
  • 结果在模拟和真实世界测试数据上均保持一致,验证了该方法在真实声学环境中的实际适用性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。