Skip to main content
QUICK REVIEW

[論文レビュー] Lightweight and perceptually-guided voice conversion for electro-laryngeal speech

Benedikt Mayrhofer, Franz Pernkopf|arXiv (Cornell University)|Jan 7, 2026
Voice and Speech Disorders被引用数 0
ひとこと要約

The paper adapts a lightweight StreamVC-based voice conversion model to electro-laryngeal (EL) speech by removing pitch/energy modules and applying self-supervised pretraining plus supervised fine-tuning with perceptual and intelligibility losses, achieving substantial CER and naturalness gains over EL, with close alignment to healthy speech.

ABSTRACT

Electro-laryngeal (EL) speech is characterized by constant pitch, limited prosody, and mechanical noise, reducing naturalness and intelligibility. We propose a lightweight adaptation of the state-of-the-art StreamVC framework to this setting by removing pitch and energy modules and combining self-supervised pretraining with supervised fine-tuning on parallel EL and healthy (HE) speech data, guided by perceptual and intelligibility losses. Objective and subjective evaluations across different loss configurations confirm their influence: the best model variant, based on WavLM features and human-feedback predictions (+WavLM+HF), drastically reduces character error rate (CER) of EL inputs, raises naturalness mean opinion score (nMOS) from 1.1 to 3.3, and consistently narrows the gap to HE ground-truth speech in all evaluated metrics. These findings demonstrate the feasibility of adapting lightweight voice conversion architectures to EL voice rehabilitation while also identifying prosody generation and intelligibility improvements as the main remaining bottlenecks.

研究の動機と目的

  • Motivate rehabilitation of electro-laryngeal speech by restoring natural prosody and reducing mechanical noise.
  • Develop a lightweight EL-to-HE voice conversion (VC) system based on adapting StreamVC architecture.
  • Investigate how perceptual and intelligibility-guided losses affect performance.
  • Provide objective and subjective evaluations across loss configurations and noisy conditions.

提案手法

  • Adapt StreamVC architecture by omitting pitch and energy modules to suit EL speech.
  • Pretrain on healthy German speech using self-supervised objectives to disentangle content and speaker features.
  • Fine-tune on parallel EL–HE data using perceptual (WavLM, WEO) and intelligibility (HF, F0, BNF, WEO) guided losses.
  • Align EL–HE data with Whisper-based DTW using Whisper Encoder Output features for robust cross-domain alignment.
  • Use GAN-based waveform synthesis with HiFi-GAN style discriminator; keep model lightweight (~30M params, 123 MB).
  • Evaluate both objective (CER, MOS predictors, SIM, F0 RMSE) and subjective metrics (nMOS, WER, SIM) under noise.

実験結果

リサーチクエスチョン

  • RQ1Can a lightweight EL-to-HE VC model achieve competitive intelligibility and naturalness compared with baselines?
  • RQ2What is the impact of perceptual/intelligibility-guided losses on EL speech restoration?
  • RQ3How close can EL-converted speech get to HE ground-truth across objective and subjective measures?
  • RQ4Is the model robust to noisy environments, and where are remaining bottlenecks (prosody, intelligibility)?

主な発見

  • Best configuration (+WavLM+HF) drastically reduces CER and increases nMOS relative to EL baseline and baselines (FreeVC, XVC).
  • Objective metrics show the proposed variant approaches HE ground-truth performance in most measures except F0 RMSE and CER.
  • WavLM or WEO perceptual losses paired with HF yield balanced improvements across intelligibility and quality; adding more losses can hurt convergence.
  • Subjective results align with objective trends, with +WavLM+HF giving the best naturalness (nMOS) and solid speaker similarity.
  • Prosody generation and intelligibility remain the main bottlenecks for EL rehabilitation; prosody improvements are highlighted as future work.
  • Under moderate-to-high SNR, the approach provides clear intelligibility gains; performance degrades more with non-stationary noise at low SNR, narrowing the gap to EL input.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。