[Paper Review] PSD estimation in Beamspace for Estimating Direct-to-Reverberant Ratio from A Reverberant Speech Signal
This paper proposes a blind DRR estimation method using power spectral density (PSD) estimation in beamspace with a microphone array, leveraging two delay-and-sum beamformers with distinct beampatterns to separate direct sound and reverberation. Evaluated on the ACE Challenge corpus, the method achieves lower estimation error variance and improved accuracy compared to prior approaches, especially under varying acoustic conditions and noise levels, with calibration further reducing bias.
A method for estimation of direct-to-reverberant ratio (DRR) using a microphone array is proposed. The proposed method estimates the power spectral density (PSD) of the direct sound and the reverberation using the algorithm extit{PSD estimation in beamspace} with a microphone array and calculates the DRR of the observed signal. The speech corpus of the ACE (Acoustic Characterisation of Environments) Challenge was utilised for evaluating the practical feasibility of the proposed method. The experimental results revealed that the proposed method was able to effectively estimate the DRR from a recording of a reverberant speech signal which included various environmental noise.
Motivation & Objective
- To develop a blind DRR estimation method that does not require prior knowledge of the room impulse response or source direction.
- To improve estimation accuracy by exploiting spatial diversity through beamforming in beamspace.
- To evaluate the method under fully blind conditions, including unknown source direction and environmental noise, using the ACE Challenge dataset.
- To reduce estimation bias caused by differences in DRR definition between the proposed method and the ACE corpus.
- To assess robustness across diverse acoustic environments and noise levels.
Proposed method
- The method uses two delay-and-sum beamformers with identical beampattern shapes but different steering directions: one aligned with the estimated source direction and another offset by 60 degrees in azimuth.
- The beamformer outputs are used to estimate the power spectral density (PSD) of the direct sound and reverberation in beamspace, based on the assumption that direct sound is spatially coherent while reverberation is diffuse.
- The DRR is calculated as the ratio of the estimated PSD of the direct sound to that of the reverberation across frequency bins.
- A voice activity detection (VAD) stage selects speech-active frames based on signal energy thresholds derived from stationary noise estimation.
- Direction-of-arrival (DOA) estimation is performed using a steered beamformer with delay-and-sum beamforming to provide the required source direction input.
- A constant bias calibration is applied using the Dev dataset to correct discrepancies in DRR definition between the method and the ACE corpus, particularly regarding early reflections.
Experimental results
Research questions
- RQ1Can a beamspace-based PSD estimation method achieve more accurate DRR estimation than coherence-based methods in a fully blind scenario?
- RQ2How does the proposed method perform across diverse acoustic environments and noise levels in the ACE Challenge dataset?
- RQ3To what extent does the method’s accuracy depend on the accuracy of the pre-estimated DOA?
- RQ4Can bias in DRR estimation be effectively compensated using a small calibration dataset?
- RQ5Does using two beamformers with identical beampatterns but different steering directions improve separation of direct and reverberant components?
Key findings
- The proposed method achieved lower variance in estimation error compared to the authors’ previous coherence-based PSD method, indicating greater robustness.
- The method showed improved accuracy across all tested noise levels and room types in the ACE Challenge’s Eval dataset.
- The bias calibration significantly reduced systematic errors, particularly due to differences in how early reflections were classified in the ACE corpus versus the proposed model.
- Performance degradation was observed in environments with poor DOA estimation, highlighting the method’s sensitivity to source direction accuracy.
- The method outperformed prior approaches in terms of both mean error and error distribution spread, especially in challenging acoustic conditions.
- The results suggest that future DRR estimation methods should be less dependent on precise DOA estimates to improve real-world robustness.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.