[Paper Review] Learning robust speech representation with an articulatory-regularized variational autoencoder
This paper proposes an articulatory-regularized variational autoencoder (AR-VAE) that improves speech representation learning by constraining part of the latent space to follow articulatory trajectories derived from EMA data. The method accelerates training convergence, reduces reconstruction loss, and enhances speech denoising performance compared to a conventional VAE, demonstrating that incorporating articulatory priors leads to more robust and efficient speech representation learning.
It is increasingly considered that human speech perception and production both rely on articulatory representations. In this paper, we investigate whether this type of representation could improve the performances of a deep generative model (here a variational autoencoder) trained to encode and decode acoustic speech features. First we develop an articulatory model able to associate articulatory parameters describing the jaw, tongue, lips and velum configurations with vocal tract shapes and spectral features. Then we incorporate these articulatory parameters into a variational autoencoder applied on spectral features by using a regularization technique that constraints part of the latent space to follow articulatory trajectories. We show that this articulatory constraint improves model training by decreasing time to convergence and reconstruction loss at convergence, and yields better performance in a speech denoising task.
Motivation & Objective
- To investigate whether prior articulatory knowledge can accelerate and improve the learning of deep speech representations in a variational autoencoder.
- To evaluate whether articulatory constraints enhance robustness of learned representations under noisy conditions.
- To develop a joint articulatory-acoustic modeling framework that integrates physiological speech production knowledge into latent space learning.
- To assess the impact of articulatory regularization on speech reconstruction quality and denoising performance.
Proposed method
- An articulatory model is trained on EMA data (jaw, tongue, lips, velum) from two reference speakers to map articulatory parameters to vocal tract shapes and spectral features.
- A variational autoencoder (VAE) is trained on spectral features (18 Bark-scale cepstral coefficients) extracted from speech audio.
- A regularization term is introduced in the VAE loss function to constrain a subset of the latent space to follow learned articulatory trajectories, enforcing alignment with articulatory dynamics.
- The regularization uses a weighted reconstruction loss between the latent code and predicted articulatory parameters, with hyperparameter α controlling the strength of the constraint.
- The model is evaluated using a conventional VAE as baseline, with training monitored for convergence speed and reconstruction loss.
- Speech denoising performance is assessed via HMM-based phonetic decoding accuracy and MUSHRA subjective quality scores across multiple SNR levels.
Experimental results
Research questions
- RQ1Can prior articulatory knowledge accelerate the training process of a VAE for speech representation learning?
- RQ2Does incorporating articulatory constraints improve the robustness of learned representations in noisy speech conditions?
- RQ3How does the AR-VAE compare to a conventional VAE in terms of reconstruction quality and convergence speed?
- RQ4Can articulatory regularization enhance the perceptual quality of reconstructed speech signals in denoising tasks?
Key findings
- The AR-VAE achieved faster convergence and lower reconstruction loss at convergence compared to the conventional VAE, with a significant reduction in training time.
- On a speech denoising task, the AR-VAE improved phonetic decoding accuracy by 12.5 percentage points (from 78.3% to 90.8%) at 0 dB SNR, with p < 0.001.
- MUSHRA scores showed a statistically significant improvement (p < 0.001) for clean speech and SNR = 10 dB, but no significant difference at SNR = 5 dB and 0 dB.
- The AR-VAE demonstrated better spectral reconstruction fidelity, particularly in preserving formant transitions and articulatory dynamics.
- The regularization with α = 1 yielded the best performance, indicating optimal balance between acoustic and articulatory constraints.
- The results suggest that articulatory priors enhance the disentanglement and interpretability of learned speech representations in the latent space.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.