[Paper Review] Does Phase Matter For Monaural Source Separation?
This paper investigates whether preserving phase information in spectral representations improves monaural source separation using a neurally plausible sparse generative model. By training a convolutional locally competitive algorithm on phase-rich and phase-discarded spectrograms, the authors show that phase preservation reduces artifacts (GSAR = 2.88 in denoised results) and achieves state-of-the-art GSIR of 19.46, demonstrating phase's critical role in high-fidelity audio separation.
The "cocktail party" problem of fully separating multiple sources from a single channel audio waveform remains unsolved. Current biological understanding of neural encoding suggests that phase information is preserved and utilized at every stage of the auditory pathway. However, current computational approaches primarily discard phase information in order to mask amplitude spectrograms of sound. In this paper, we seek to address whether preserving phase information in spectral representations of sound provides better results in monaural separation of vocals from a musical track by using a neurally plausible sparse generative model. Our results demonstrate that preserving phase information reduces artifacts in the separated tracks, as quantified by the signal to artifact ratio (GSAR). Furthermore, our proposed method achieves state-of-the-art performance for source separation, as quantified by a mean signal to interference ratio (GSIR) of 19.46.
Motivation & Objective
- To investigate whether phase information in spectral representations enhances monaural source separation performance.
- To evaluate the role of phase in neural-inspired sparse coding models for audio separation.
- To compare phase-preserving vs. phase-discarded approaches in separating vocals from music using a generative sparse model.
- To achieve state-of-the-art performance in monaural source separation while maintaining biological plausibility.
Proposed method
- Trained a convolutional locally competitive algorithm (LCA) on 2-second audio clips from the MIR-1K dataset using phase-rich and phase-discarded spectrograms.
- Used a soft-thresholding activation function with membrane potentials to enforce sparsity and neural plausibility.
- Applied iterative stochastic gradient descent with a local Hebbian rule and momentum to optimize basis vectors.
- Optimized the sparsity threshold via denoising performance, finding 0.625 as optimal, yielding ~2.8% average sparsity.
- Generated sparse codes for training and test sets, then used a linear classifier trained on these codes to separate vocals and non-vocals.
- Applied a second denoising pass on separated stems to refine results, achieving the best performance.
Experimental results
Research questions
- RQ1Does preserving phase information in spectral representations lead to fewer artifacts in monaural source separation?
- RQ2Can a neurally plausible sparse generative model achieve state-of-the-art performance in vocal separation from music?
- RQ3How does phase preservation affect signal-to-interference ratio (SIR) and signal-to-artifact ratio (SAR) in blind source separation?
- RQ4Is phase information necessary for learning accurate spectrotemporal receptive fields in a sparse coding framework?
Key findings
- Preserving phase information significantly reduces artifacts, as evidenced by a GSAR of 2.88 in the denoised results, compared to -9.72 for phase-discarded models.
- The phase-preserving model achieved a mean GSIR of 19.46 in the denoised condition, representing state-of-the-art performance for monaural source separation.
- Despite slightly lower GSIR values in non-denoised phase results (12.12), the phase model consistently showed better artifact suppression than phase-discarded alternatives.
- The denoised phase-based separation produced a GNSDR of 5.21, indicating improved overall source quality and reduced distortion.
- The optimal sparsity threshold was found to be 0.625, yielding approximately 2.8% average sparsity across the dataset.
- Visual analysis confirmed that phase-preserving reconstruction preserved spectral structure better and avoided spurious energy in high-frequency regions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.