[Paper Review] Phase-aware Speech Enhancement with Deep Complex U-Net
This paper introduces a Deep Complex U-Net for phase-aware speech enhancement, a polar complex masking scheme, and a wSDR loss to improve reconstruction quality beyond magnitude-only methods.
Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction. This is due to the difficulty of estimating the phase of clean speech. To improve speech enhancement performance, we tackle the phase estimation problem in three ways. First, we propose Deep Complex U-Net, an advanced U-Net structured model incorporating well-defined complex-valued building blocks to deal with complex-valued spectrograms. Second, we propose a polar coordinate-wise complex-valued masking method to reflect the distribution of complex ideal ratio masks. Third, we define a novel loss function, weighted source-to-distortion ratio (wSDR) loss, which is designed to directly correlate with a quantitative evaluation measure. Our model was evaluated on a mixture of the Voice Bank corpus and DEMAND database, which has been widely used by many deep learning models for speech enhancement. Ablation experiments were conducted on the mixed dataset showing that all three proposed approaches are empirically valid. Experimental results show that the proposed method achieves state-of-the-art performance in all metrics, outperforming previous approaches by a large margin.
Motivation & Objective
- Motivate improved speech enhancement by addressing phase estimation beyond noisy-phase reuse.
- Develop a Deep Complex U-Net with complex-valued building blocks for complex spectrograms.
- Propose a polar coordinate-wise complex-valued masking method to better reflect complex masks distribution.
- Introduce a weighted source-to-distortion ratio (wSDR) loss aligned with evaluation metrics.
- Demonstrate empirical gains via ablation on a standard mixed speech dataset.
Proposed method
- Extend U-Net with complex-valued layers to operate on complex spectrograms.
- Introduce polar coordinate-wise complex-valued masking to model phase and magnitude jointly.
- Define and employ a weighted SDR (wSDR) loss to correlate with quantitative metrics.
- Evaluate on the Voice Bank + DEMAND mixed dataset and perform ablation studies.
- Compare to prior methods that estimate magnitude and reuse noisy phase.
Experimental results
Research questions
- RQ1Can a complex-valued U-Net improve phase-aware speech enhancement over magnitude-focused models?
- RQ2Does polar coordinate-wise complex masking better capture the distribution of complex masks than real-valued masking?
- RQ3Does the wSDR loss directly improve alignment with objective evaluation metrics?
- RQ4What is the contribution of each proposed component (complex U-Net, polar masking, wSDR) to overall performance?
- RQ5How does the proposed method perform on standard mixed speech datasets compared to prior approaches?
Key findings
- The approach achieves state-of-the-art performance in all metrics on the mixed Voice Bank and DEMAND dataset.
- Ablation experiments confirm empirical validity of all three proposed approaches.
- The model outperforms previous approaches by a large margin according to the abstract.
- The combination of complex-valued modeling, polar masking, and wSDR loss yields improved enhancement results.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.