Skip to main content
QUICK REVIEW

[Paper Review] Evaluation of a Multi-Resolution Dyadic Wavelet Transform Method for usable Speech Detection

Wajdi Ghezaiel, Amel Ben Slimane Rahmouni|arXiv (Cornell University)|Jan 2, 2013
Speech and Audio Processing3 citations
TL;DR

This paper proposes a multi-resolution dyadic wavelet transform (MRDWT) method to detect usable speech segments in co-channel speech mixtures, leveraging time-frequency analysis to isolate speech components. Evaluated on the TIMIT database, the method achieves 95.76% usable speech detection accuracy with 29.65% false alarms across diverse gender speaker mixtures.

ABSTRACT

Many applications of speech communication and speaker identification suffer from the problem of co-channel speech. This paper deals with a multi-resolution dyadic wavelet transform method for usable segments of co-channel speech detection that could be processed by a speaker identification system. Evaluation of this method is performed on TIMIT database referring to the Target to Interferer Ratio measure. Co-channel speech is constructed by mixing all possible gender speakers. Results do not show much difference for different mixtures. For the overall mixtures 95.76% of usable speech is correctly detected with false alarms of 29.65%.

Motivation & Objective

  • Address the challenge of co-channel speech degradation in speaker identification systems.
  • Improve the reliability of speech processing in overlapping speech environments.
  • Develop a robust method to detect usable speech segments in mixed signals without prior speaker separation.
  • Evaluate performance across diverse gender speaker combinations to ensure generalization.
  • Assess detection accuracy and false alarm rates using the Target-to-Interferer Ratio (TIR) metric.

Proposed method

  • Apply a multi-resolution dyadic wavelet transform (MRDWT) to decompose mixed speech signals into multiple frequency subbands.
  • Use time-frequency representation to identify energy-dominant regions corresponding to active speech segments.
  • Leverage the dyadic wavelet structure to achieve multirate analysis with logarithmic frequency resolution.
  • Detect usable speech by thresholding wavelet coefficients based on energy concentration and temporal continuity.
  • Utilize the TIMIT database for training and evaluation, simulating all possible gender speaker mixtures.
  • Apply the Target-to-Interferer Ratio (TIR) as a performance metric to quantify detection quality.

Experimental results

Research questions

  • RQ1Can the MRDWT method effectively detect usable speech segments in co-channel speech mixtures?
  • RQ2How does detection performance vary across different gender speaker combinations in mixed signals?
  • RQ3What is the trade-off between usable speech detection accuracy and false alarm rate in overlapping speech?
  • RQ4To what extent does the multi-resolution nature of the dyadic wavelet transform enhance segment detection compared to single-resolution methods?
  • RQ5How does the method perform under realistic, diverse speaker mixture conditions?

Key findings

  • The MRDWT method achieved a usable speech detection rate of 95.76% across all tested co-channel speech mixtures.
  • The method exhibited a false alarm rate of 29.65%, indicating moderate sensitivity to non-speech or interferer-dominant segments.
  • Performance remained consistent across different gender speaker combinations, with no significant variation in detection accuracy.
  • The TIMIT database was successfully used to simulate realistic co-channel speech conditions for evaluation.
  • The multi-resolution dyadic wavelet transform effectively captured time-frequency energy patterns associated with speech activity.
  • The method demonstrated robustness in identifying usable speech segments despite the absence of prior speaker separation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.