[Paper Review] Design and Optimization of a Speech Recognition Front-End for Distant-Talking Control of a Music Playback Device
This paper proposes a robust single-microphone speech recognition front-end for distant-talking control of portable music devices, combining cascaded echo cancellation, double-talk detection, and a novel adaptive quasi-binary mask for noise suppression. Optimized via genetic algorithm to maximize recognition accuracy, the system achieves over 90% command recognition at speech-to-music ratios as low as -35 dB, enabling reliable voice control in highly adverse acoustic conditions.
This paper addresses the challenging scenario for the distant-talking control of a music playback device, a common portable speaker with four small loudspeakers in close proximity to one microphone. The user controls the device through voice, where the speech-to-music ratio can be as low as -30 dB during music playback. We propose a speech enhancement front-end that relies on known robust methods for echo cancellation, double-talk detection, and noise suppression, as well as a novel adaptive quasi-binary mask that is well suited for speech recognition. The optimization of the system is then formulated as a large scale nonlinear programming problem where the recognition rate is maximized and the optimal values for the system parameters are found through a genetic algorithm. We validate our methodology by testing over the TIMIT database for different music playback levels and noise types. Finally, we show that the proposed front-end allows a natural interaction with the device for limited-vocabulary voice commands.
Motivation & Objective
- To address the challenge of distant-talking voice control in portable music devices with high music playback levels and low speech-to-echo ratios.
- To design a single-microphone front-end that enables reliable speech recognition under extreme acoustic degradation, including music playback and background noise.
- To optimize system parameters for maximum automatic speech recognition (ASR) performance using a nonlinear programming approach.
- To validate the system on real-world recordings with natural voice commands at varying speech-to-music ratios.
- To demonstrate feasibility of limited-vocabulary voice control in practical, real-time scenarios with low signal-to-noise ratios.
Proposed method
- The front-end employs a cascaded structure of two robust acoustic echo cancelers (RAEC) with multi-delay adaptive filtering to suppress echo from loudspeakers.
- It uses a double-talk probability (DTP) estimator and coherence-based residual echo power estimator to detect and suppress residual echo.
- A novel adaptive quasi-binary mask combines minimum mean-squared error (MMSE) estimation with ideal binary mask (IBM) approximation for enhanced noise suppression at low SNRs.
- The system integrates a voice activity detector (VAD) using a priori and posteriori SNR-based energy thresholding to isolate speech frames.
- Parameter tuning is formulated as a large-scale nonlinear programming problem, solved via a genetic algorithm (GA) to maximize ASR accuracy on the TIMIT database.
- The training data is synthesized by convolving clean TIMIT speech with room impulse responses and mixing with music, babble, or factory noise at varying levels.
Experimental results
Research questions
- RQ1Can a single-microphone front-end achieve robust speech recognition in portable music devices with high music playback levels and low speech-to-echo ratios?
- RQ2How effective is a cascaded RAEC with double-talk detection and residual echo suppression in reducing echo and noise under real-time playback conditions?
- RQ3Does combining MMSE-based enhancement with an adaptive quasi-binary mask improve recognition performance at very low SNRs compared to conventional methods?
- RQ4Can genetic algorithm-based optimization of front-end parameters significantly improve ASR accuracy compared to quality-based optimization (e.g., POLQA)?
- RQ5What is the achievable recognition rate for limited-vocabulary voice commands in real-world conditions with speech-to-music ratios as low as -35 dB?
Key findings
- The proposed front-end achieved a phone accuracy of 37.4% (bigram) on the noisy TIMIT database under babble noise, significantly outperforming baseline methods.
- In real-world testing, the system achieved 90% command recognition accuracy for 'PLAY' and 'NEXT' at speech-to-music ratios of -25 to -20 dB, and 80% for 'PLAY' at -30 to -25 dB.
- At the lowest SER of -35 to -30 dB, the system achieved 73% recognition for 'BACK' and 76% for 'PAUSE', outperforming the 25% baseline when no enhancement was applied.
- The ASR-optimized parameters outperformed POLQA-optimized parameters across all SER levels, indicating that ASR-focused tuning is more effective than quality-focused tuning for recognition tasks.
- The system preserved speech intelligibility and structure, as shown in spectrograms, with processed signals becoming clearly recognizable to human listeners despite initial inaudibility.
- The estimated speech-to-echo ratio (SER) in real experiments ranged from -35 to -20 dB, validating the generalization of the simulation-based tuning methodology to real-world conditions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.