Skip to main content
QUICK REVIEW

[Paper Review] Design and Optimization of a Speech Recognition Front-End for Distant-Talking Control of a Music Playback Device

Ramin Pichevar, Jason Wung|arXiv (Cornell University)|May 5, 2014
Speech and Audio Processing37 references3 citations
TL;DR

This paper proposes a robust single-microphone speech recognition front-end for distant-talking control of portable music devices, combining cascaded echo cancellation, double-talk detection, and a novel adaptive quasi-binary mask for noise suppression. Optimized via genetic algorithm to maximize recognition accuracy, the system achieves over 90% command recognition at speech-to-music ratios as low as -35 dB, enabling reliable voice control in highly adverse acoustic conditions.

ABSTRACT

This paper addresses the challenging scenario for the distant-talking control of a music playback device, a common portable speaker with four small loudspeakers in close proximity to one microphone. The user controls the device through voice, where the speech-to-music ratio can be as low as -30 dB during music playback. We propose a speech enhancement front-end that relies on known robust methods for echo cancellation, double-talk detection, and noise suppression, as well as a novel adaptive quasi-binary mask that is well suited for speech recognition. The optimization of the system is then formulated as a large scale nonlinear programming problem where the recognition rate is maximized and the optimal values for the system parameters are found through a genetic algorithm. We validate our methodology by testing over the TIMIT database for different music playback levels and noise types. Finally, we show that the proposed front-end allows a natural interaction with the device for limited-vocabulary voice commands.

Motivation & Objective

  • To address the challenge of distant-talking voice control in portable music devices with high music playback levels and low speech-to-echo ratios.
  • To design a single-microphone front-end that enables reliable speech recognition under extreme acoustic degradation, including music playback and background noise.
  • To optimize system parameters for maximum automatic speech recognition (ASR) performance using a nonlinear programming approach.
  • To validate the system on real-world recordings with natural voice commands at varying speech-to-music ratios.
  • To demonstrate feasibility of limited-vocabulary voice control in practical, real-time scenarios with low signal-to-noise ratios.

Proposed method

  • The front-end employs a cascaded structure of two robust acoustic echo cancelers (RAEC) with multi-delay adaptive filtering to suppress echo from loudspeakers.
  • It uses a double-talk probability (DTP) estimator and coherence-based residual echo power estimator to detect and suppress residual echo.
  • A novel adaptive quasi-binary mask combines minimum mean-squared error (MMSE) estimation with ideal binary mask (IBM) approximation for enhanced noise suppression at low SNRs.
  • The system integrates a voice activity detector (VAD) using a priori and posteriori SNR-based energy thresholding to isolate speech frames.
  • Parameter tuning is formulated as a large-scale nonlinear programming problem, solved via a genetic algorithm (GA) to maximize ASR accuracy on the TIMIT database.
  • The training data is synthesized by convolving clean TIMIT speech with room impulse responses and mixing with music, babble, or factory noise at varying levels.

Experimental results

Research questions

  • RQ1Can a single-microphone front-end achieve robust speech recognition in portable music devices with high music playback levels and low speech-to-echo ratios?
  • RQ2How effective is a cascaded RAEC with double-talk detection and residual echo suppression in reducing echo and noise under real-time playback conditions?
  • RQ3Does combining MMSE-based enhancement with an adaptive quasi-binary mask improve recognition performance at very low SNRs compared to conventional methods?
  • RQ4Can genetic algorithm-based optimization of front-end parameters significantly improve ASR accuracy compared to quality-based optimization (e.g., POLQA)?
  • RQ5What is the achievable recognition rate for limited-vocabulary voice commands in real-world conditions with speech-to-music ratios as low as -35 dB?

Key findings

  • The proposed front-end achieved a phone accuracy of 37.4% (bigram) on the noisy TIMIT database under babble noise, significantly outperforming baseline methods.
  • In real-world testing, the system achieved 90% command recognition accuracy for 'PLAY' and 'NEXT' at speech-to-music ratios of -25 to -20 dB, and 80% for 'PLAY' at -30 to -25 dB.
  • At the lowest SER of -35 to -30 dB, the system achieved 73% recognition for 'BACK' and 76% for 'PAUSE', outperforming the 25% baseline when no enhancement was applied.
  • The ASR-optimized parameters outperformed POLQA-optimized parameters across all SER levels, indicating that ASR-focused tuning is more effective than quality-focused tuning for recognition tasks.
  • The system preserved speech intelligibility and structure, as shown in spectrograms, with processed signals becoming clearly recognizable to human listeners despite initial inaudibility.
  • The estimated speech-to-echo ratio (SER) in real experiments ranged from -35 to -20 dB, validating the generalization of the simulation-based tuning methodology to real-world conditions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.