Skip to main content

Eun-Woo Song

Yonsei University · Computer Science

About the Lab

Professor Eun-Woo Song's research lab specializes in neural speech synthesis and vocoding, focusing on developing high-fidelity, efficient, and compact waveform generation techniques for text-to-speech systems. The lab explores deep learning-based approaches—particularly generative adversarial networks (GANs) and recurrent neural networks (RNNs)—to model speech excitation and spectral envelopes with improved perceptual quality. Key research directions include distillation-free neural vocoders, time-frequency trajectory modeling, and perceptual loss optimization for faster and more natural speech synthesis.

neural vocodertext-to-speechwaveform generationspeech synthesisadversarial training

Research Overview

Papers
86
Total Citations
1,569
Papers (5y)
21
Primary Field
Computer Science

Research Output Trend

Figures are computed from collected data and may differ slightly.

Publications per year (5y)
21total
2022
2023
2024
2025
2026
Citations per year (5y)
78total
20222023202420252026

Selected Papers

15
1
Article|804 citations·2020
Parallel Wavegan: A Fast Waveform Generation Model Based on Generative Adversarial Networks with Multi-Resolution Spectrogram
Ryuichi Yamamoto, Eunwoo Song, Jae-Min Kim

We propose Parallel WaveGAN, a distillation-free, fast, and small-footprint waveform generation method using a generative adversarial network. In the proposed method, a non-autoregressive WaveNet is trained by jointly optimizing multi-resolution spectrogram and adversarial loss functions, which can effectively capture the time-frequency distribution of the realistic speech waveform. As our method does not require density distillation used in the conventional teacher-student framework, the entire

Signal ProcessingComputer Science
2
Article|73 citations·2017
Effective Spectral and Excitation Modeling Techniques for LSTM-RNN-Based Speech Synthesis Systems
Eunwoo Song, Frank K. Soong, Hong-Goo Kang
SJR Q1IEEE/ACM Transactions on Audio Speech and Language Processing

In this paper, we report research results on modeling the parameters of an improved time-frequency trajectory excitation (ITFTE) and spectral envelopes of an LPC vocoder with a long short-term memory (LSTM)-based recurrent neural network (RNN) for high-quality text-to-speech (TTS) systems. The ITFTE vocoder has been shown to significantly improve the perceptual quality of statistical parameter-based TTS systems in our prior works. However, a simple feed-forward deep neural network (DNN) with a f

Artificial IntelligenceComputer Science
3
Article|38 citations·2019
ExcitNet Vocoder: A Neural Excitation Model for Parametric Speech Synthesis Systems
Eunwoo Song, Kyungguen Byun, Hong-Goo Kang

This paper proposes a WaveNet-based neural excitation model (ExcitNet) for statistical parametric speech synthesis systems. Conventional WaveNet-based neural vocoding systems significantly improve the perceptual quality of synthesized speech by statistically generating a time sequence of speech waveforms through an auto-regressive framework. However, they often suffer from noisy outputs because of the difficulties in capturing the complicated time-varying nature of speech signals. To improve mod

Artificial IntelligenceComputer Science
4
Article|13 citations·2021
Improved Parallel Wavegan Vocoder with Perceptually Weighted Spectrogram Loss
Eunwoo Song, Ryuichi Yamamoto, Min-Jae Hwang, Jinseob Kim, Ohsung Kwon, Jae-Min Kim

This paper proposes a spectral-domain perceptual weighting technique for Parallel WaveGAN-based text-to-speech (TTS) systems. The recently proposed Parallel WaveGAN vocoder successfully generates waveform sequences using a fast non-autoregressive WaveNet model. By employing multi-resolution short-time Fourier transform (MR-STFT) criteria with a generative adversarial network, the light-weight convolutional networks can be effectively trained without any distillation process. To further improve t

Signal ProcessingComputer Science
5
Article|9 citations·2015
Improved time-frequency trajectory excitation modeling for a statistical parametric speech synthesis system
Eunwoo Song, Young-Sun Joo, Hong-Goo Kang

This paper proposes an improved time-frequency trajectory excitation (TFTE) modeling method for a statistical parametric speech synthesis system. The proposed approach overcomes the dimensional variation problem of the training process caused by the inherent nature of the pitch-dependent analysis paradigm. By reducing the redundancies of the parameters using predicted average block coefficients (PABC), the proposed algorithm efficiently models excitation, even if its dimension is varied. Objecti

Artificial IntelligenceComputer Science
6
Article|8 citations·2015
Deep neural network-based statistical parametric speech synthesis system using improved time-frequency trajectory excitation model
Eunwoo Song, Hong-Goo Kang
Artificial IntelligenceComputer Science
7
Article|7 citations·2020
Neural Text-to-Speech with a Modeling-by-Generation Excitation Vocoder
Eunwoo Song, Min-Jae Hwang, Ryuichi Yamamoto, Jinseob Kim, Ohsung Kwon, Jae-Min Kim

This paper proposes a modeling-by-generation (MbG) excitation vocoder for a neural text-to-speech (TTS) system. Recently proposed neural excitation vocoders can realize qualified waveform generation by combining a vocal tract filter with a WaveNet-based glottal excitation generator. However, when these vocoders are used in a TTS system, the quality of synthesized speech is often degraded owing to a mismatch between training and synthesis steps. Specifically, the vocoder is separately trained fro

Artificial IntelligenceComputer Science
8
Article|5 citations·2022
TTS-by-TTS 2: Data-Selective Augmentation for Neural Speech Synthesis Using Ranking Support Vector Machine with Variational Autoencoder
Eunwoo Song, Ryuichi Yamamoto, Ohsung Kwon, Chanho Song, Min-Jae Hwang, Suhyeon Oh, Hyun-Wook Yoon, Jinseob Kim, Jae-Min Kim
Interspeech 2022

Recent advances in synthetic speech quality have enabled us to train text-to-speech (TTS) systems by using synthetic corpora. However, merely increasing the amount of synthetic data is not always advantageous for improving training efficiency. Our aim in this study is to selectively choose synthetic data that are beneficial to the training process. In the proposed method, we first adopt a variational autoencoder whose posterior distribution is utilized to extract latent features representing aco

Artificial IntelligenceComputer Science
9
Preprint|4 citations·2024
Improved Parallel Wavegan Vocoder With Perceptually Weighted Spectrogram Loss
Eunwoo Song, Ryuichi Yamamoto, Min-Jae Hwang, Jinseob Kim, Ohsung Kwon, Jae-Min Kim
IEEE RESOURCE CENTERSOA

This paper proposes a spectral-domain perceptual weighting technique for Parallel WaveGAN-based text-to-speech (TTS) systems. The recently proposed Parallel WaveGAN vocoder successfully generates waveform sequences using a fast non-autoregressive WaveNet model. By employing multi-resolution short-time Fourier transform (MR-STFT) criteria with a generative adversarial network, the light-weight convolutional networks can be effectively trained without any distillation process. To further improve t

Signal ProcessingComputer Science
10
Article|4 citations·2016
Improved Time-Frequency Trajectory Excitation Vocoder for DNN-Based Speech Synthesis
Eunwoo Song, Frank K. Soong, Hong-Goo Kang
Artificial IntelligenceComputer Science
11
Article|3 citations·2020
Speaker-Adaptive Neural Vocoders for Parametric Speech Synthesis Systems
Eunwoo Song, Jinseob Kim, Kyungguen Byun, Hong-Goo Kang

This paper proposes speaker-adaptive neural vocoders for parametric text-to-speech (TTS) systems. Recently proposed WaveNet-based neural vocoding systems successfully generate a time sequence of speech signal with an autoregressive framework. However, it remains a challenge to synthesize high-quality speech when the amount of a target speaker's training data is insufficient. To generate more natural speech signals with the constraint of limited training data, we propose a speaker adaptation task

Artificial IntelligenceComputer Science
12
Preprint|3 citations·2021
Improved parallel WaveGAN vocoder with perceptually weighted spectrogram loss
Eunwoo Song, Ryuichi Yamamoto, Min-Jae Hwang, Jinseob Kim, Ohsung Kwon, Jae-Min Kim
arXiv (Cornell University)OA

This paper proposes a spectral-domain perceptual weighting technique for Parallel WaveGAN-based text-to-speech (TTS) systems. The recently proposed Parallel WaveGAN vocoder successfully generates waveform sequences using a fast non-autoregressive WaveNet model. By employing multi-resolution short-time Fourier transform (MR-STFT) criteria with a generative adversarial network, the light-weight convolutional networks can be effectively trained without any distillation process. To further improve t

Signal ProcessingComputer Science
13
Article|2 citations·2017
Perceptual quality and modeling accuracy of excitation parameters in DLSTM-based speech synthesis systems
Eunwoo Song, Frank K. Soong, Hong-Goo Kang
2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)

This paper investigates how the perceptual quality of the synthesized speech is affected by reconstruction errors in excitation signals generated by a deep learning-based statistical model. In this framework, the excitation signal obtained by an LPC inverse filter is first decomposed into harmonic and noise components using an improved time-frequency trajectory excitation (ITFTE) scheme, then they are trained and generated by a deep long short-term memory (DLSTM)-based speech synthesis system. B

Artificial IntelligenceComputer Science
14
Preprint|2 citations·2018
ExcitNet vocoder: A neural excitation model for parametric speech synthesis systems
Eunwoo Song, Kyungguen Byun, Hong-Goo Kang
arXiv (Cornell University)OA

This paper proposes a WaveNet-based neural excitation model (ExcitNet) for statistical parametric speech synthesis systems. Conventional WaveNet-based neural vocoding systems significantly improve the perceptual quality of synthesized speech by statistically generating a time sequence of speech waveforms through an auto-regressive framework. However, they often suffer from noisy outputs because of the difficulties in capturing the complicated time-varying nature of speech signals. To improve mod

Artificial IntelligenceComputer Science
15
Article|2 citations·2016
Multi-class learning algorithm for deep neural network-based statistical parametric speech synthesis
Eunwoo Song, Hong-Goo Kang

This paper proposes a multi-class learning (MCL) algorithm for a deep neural network (DNN)-based statistical parametric speech synthesis (SPSS) system. Although the DNN-based SPSS system improves the modeling accuracy of statistical parameters, its synthesized speech is often muffled because the training process only considers the global characteristics of the entire set of training data, but does not explicitly consider any local variations. We introduce a DNN-based context clustering algorithm

Artificial IntelligenceComputer Science

Research Areas

Artificial IntelligenceSignal ProcessingMolecular BiologyElectrical and Electronic EngineeringObstetrics and GynecologyBiomedical Engineering

Dive deeper into Eun-Woo Song's research on Nubint

Open this lab's papers in the app to read with AI, summarize, and cite in your writing.