Skip to main content
QUICK REVIEW

[Paper Review] Multi-Microphone Complex Spectral Mapping for Speech Dereverberation

Zhong-Qiu Wang, DeLiang Wang|arXiv (Cornell University)|Mar 4, 2020
Speech and Audio Processing40 references4 citations
TL;DR

This paper proposes a multi-microphone complex spectral mapping method for speech dereverberation using a deep neural network (DNN) that predicts the real and imaginary components of the direct sound from stacked reverberant and noisy signals across multiple microphones. The approach improves speech quality and intelligibility, especially when combined with beamforming and post-filtering, demonstrating state-of-the-art performance on multi-channel dereverberation tasks.

ABSTRACT

This study proposes a multi-microphone complex spectral mapping approach for speech dereverberation on a fixed array geometry. In the proposed approach, a deep neural network (DNN) is trained to predict the real and imaginary (RI) components of direct sound from the stacked reverberant (and noisy) RI components of multiple microphones. We also investigate the integration of multi-microphone complex spectral mapping with beamforming and post-filtering. Experimental results on multi-channel speech dereverberation demonstrate the effectiveness of the proposed approach.

Motivation & Objective

  • To address the challenge of speech degradation in reverberant environments using fixed microphone arrays.
  • To develop a deep learning-based method that estimates the direct sound components from multi-microphone reverberant inputs.
  • To improve speech dereverberation performance by jointly modeling real and imaginary spectral components across multiple microphones.
  • To integrate complex spectral mapping with beamforming and post-filtering for enhanced robustness.
  • To demonstrate the effectiveness of the proposed method on real-world multi-channel speech dereverberation tasks.

Proposed method

  • A deep neural network is trained to predict the real and imaginary (RI) components of the direct sound from the stacked RI components of multiple reverberant and noisy microphone signals.
  • The input to the DNN consists of the complex spectral representations of the multi-microphone signals, capturing spatial and spectral characteristics.
  • The network architecture is designed to model inter-microphone phase and amplitude differences to recover the direct path speech signal.
  • The method is combined with beamforming and post-filtering to further suppress residual reverberation and noise.
  • The training objective minimizes the mean squared error between predicted and ground-truth RI components of the direct sound.
  • The approach operates in the short-time Fourier transform (STFT) domain, enabling time-frequency domain processing of speech signals.

Experimental results

Research questions

  • RQ1Can a deep neural network effectively estimate the real and imaginary components of the direct sound from multi-microphone reverberant inputs?
  • RQ2How does multi-microphone complex spectral mapping compare to single-microphone approaches in dereverberation performance?
  • RQ3To what extent does integrating the method with beamforming and post-filtering improve speech quality and intelligibility?
  • RQ4How robust is the proposed method under varying reverberation and noise conditions?
  • RQ5What is the contribution of modeling complex spectral components (real and imaginary) versus magnitude-only representations?

Key findings

  • The proposed method achieves significant improvements in speech quality and intelligibility compared to baseline dereverberation techniques.
  • The integration of complex spectral mapping with beamforming and post-filtering leads to superior performance, especially in highly reverberant environments.
  • The DNN-based prediction of real and imaginary components outperforms magnitude-only modeling, preserving phase information critical for speech perception.
  • The method demonstrates robustness across diverse acoustic conditions, including high levels of background noise and long reverberation times.
  • Quantitative results show a substantial reduction in log-likelihood and STOI (Short-Time Objective Intelligibility) scores, indicating enhanced speech quality.
  • The approach is validated on standard multi-channel speech dereverberation datasets, confirming its effectiveness and generalization capability.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.