[Paper Review] Bridging the Gap Between Monaural Speech Enhancement and Recognition with Distortion-Independent Acoustic Modeling
This paper proposes distortion-independent acoustic modeling to bridge the gap between monaural speech enhancement and automatic speech recognition (ASR), where enhancement-induced distortions typically degrade ASR performance. By training acoustic models on diverse distortions (including noise, reverberation, and enhancement artifacts), the method generalizes to unseen distortions, achieving state-of-the-art WER of 9.2% on CHiME-2, outperforming prior systems by 6.5% relatively and demonstrating robustness across multiple enhancement frontends.
Monaural speech enhancement has made dramatic advances since the introduction of deep learning a few years ago. Although enhanced speech has been demonstrated to have better intelligibility and quality for human listeners, feeding it directly to automatic speech recognition (ASR) systems trained with noisy speech has not produced expected improvements in ASR performance. The lack of an enhancement benefit on recognition, or the gap between monaural speech enhancement and recognition, is often attributed to speech distortions introduced in the enhancement process. In this study, we analyze the distortion problem, compare different acoustic models, and investigate a distortion-independent training scheme for monaural speech recognition. Experimental results suggest that distortion-independent acoustic modeling is able to overcome the distortion problem. Such an acoustic model can also work with speech enhancement models different from the one used during training. Moreover, the models investigated in this paper outperform the previous best system on the CHiME-2 corpus.
Motivation & Objective
- To address the persistent performance gap between monaural speech enhancement and ASR, where enhanced speech often fails to improve recognition due to distortion.
- To investigate whether acoustic models trained to be invariant to various distortions can generalize to untrained enhancement artifacts.
- To evaluate the effectiveness of distortion-independent acoustic modeling across different speech enhancement frontends and noise conditions.
- To surpass the performance of prior systems on the CHiME-2 benchmark, particularly in reverberant and noisy environments.
- To demonstrate that large-scale training on diverse distortions enables generalization to unseen distortions, including those from different enhancement models.
Proposed method
- The study trains acoustic models on a large-scale dataset of speech corrupted by 10,000 diverse noise types, including additive noise, reverberation (RIRs), and enhancement artifacts.
- It introduces a distortion-independent acoustic modeling approach that treats all distortions as part of the training distribution, avoiding reliance on clean speech or specific enhancement outputs.
- The method uses utterance-wise recurrent dropout during training to improve generalization and robustness to sequence-level variations.
- It compares five acoustic model types: noise-independent, reverberation-independent, distortion-independent, and others, under varying enhancement conditions.
- The model is evaluated with multiple enhancement frontends (LSTM, CRN, IRM) on the CHiME-2 and ADT datasets, using WER as the primary metric.
- The training strategy explicitly avoids speaker adaptation to isolate the effect of distortion robustness on recognition performance.
Experimental results
Research questions
- RQ1Can acoustic models trained to be invariant to a wide range of distortions generalize to enhancement-induced distortions not seen during training?
- RQ2Does distortion-independent acoustic modeling outperform traditional noise- or reverberation-dependent models in ASR when used with enhanced speech?
- RQ3Can a single distortion-independent acoustic model effectively work with multiple speech enhancement frontends, even when the enhancement model differs from the one used during training?
- RQ4To what extent does large-scale training on diverse distortions improve generalization to untrained distortions in real-world ASR scenarios?
- RQ5How does the performance of distortion-independent modeling compare to prior state-of-the-art systems on the CHiME-2 benchmark?
Key findings
- The distortion-independent acoustic model achieves a 9.2% average WER on the CHiME-2 corpus, outperforming the previous best system by 6.5% relatively.
- The noise-dependent acoustic model achieves an average WER of 8.7%, also surpassing prior state-of-the-art systems.
- On 9dB and 3dB SNR conditions, enhanced speech using the distortion-independent model yields lower WER than unenhanced speech, confirming the benefit of enhancement when combined with robust acoustic modeling.
- The distortion-independent model performs well across multiple enhancement frontends (LSTM, CRN, IRM), indicating generalization to different types of enhancement artifacts.
- The model generalizes to untrained distortions, including RIR mismatches and unseen enhancement methods, demonstrating strong robustness.
- Even with 3.4% average WER on reverberant speech, the model shows that distortion-invariant training enables effective recognition in challenging acoustic conditions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.