Skip to main content

송은우 교수

Eun-Woo Song

연세대학교 전기전자공학부 · 컴퓨터과학

연구실 소개

송은우 교수의 연구실은 신경망 기반 음성 합성 및 음성 생성 기술에 중점을 두고 있으며, 특히 비자율적(Non-autoregressive)이고 경량화된 웨이브포맷 생성 모델인 Parallel WaveGAN을 핵심으로 연구를 전개하고 있습니다. 고해상도 음성 품질을 달성하기 위해 주파수 도메인에서의 인지적 중요도를 반영한 손실 함수 설계와 함께, 음성의 염기 신호(Excitation)를 정밀하게 모델링하는 기술적 접근도 함께 연구하고 있습니다. 특히, 음성 합성의 전반적 품질 향상을 위해 음성 파arameter 추정과 음성 생성 간의 일관성 문제를 해결하기 위한 통합적 모델링 기법 개발에도 기여하고 있습니다.

Parallel WaveGAN음성 합성신경 음성 생성비자율 음성 생성excitation 모델링

연구 현황

논문 수
86
총 인용 수
1,569
최근 5년 논문
21
주요 분야
컴퓨터과학

연구 성과 추이

표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.

5개년 연도별 논문 게재 수
21총합
2022
2023
2024
2025
2026
5개년 연도별 피인용 수
78총합
20222023202420252026

주요 논문

15
1
논문|인용수 804·2020
Parallel Wavegan: A Fast Waveform Generation Model Based on Generative Adversarial Networks with Multi-Resolution Spectrogram
Ryuichi Yamamoto, Eunwoo Song, Jae-Min Kim

We propose Parallel WaveGAN, a distillation-free, fast, and small-footprint waveform generation method using a generative adversarial network. In the proposed method, a non-autoregressive WaveNet is trained by jointly optimizing multi-resolution spectrogram and adversarial loss functions, which can effectively capture the time-frequency distribution of the realistic speech waveform. As our method does not require density distillation used in the conventional teacher-student framework, the entire

Signal ProcessingComputer Science
2
논문|인용수 73·2017
Effective Spectral and Excitation Modeling Techniques for LSTM-RNN-Based Speech Synthesis Systems
Eunwoo Song, Frank K. Soong, Hong-Goo Kang
SJR Q1IEEE/ACM Transactions on Audio Speech and Language Processing

In this paper, we report research results on modeling the parameters of an improved time-frequency trajectory excitation (ITFTE) and spectral envelopes of an LPC vocoder with a long short-term memory (LSTM)-based recurrent neural network (RNN) for high-quality text-to-speech (TTS) systems. The ITFTE vocoder has been shown to significantly improve the perceptual quality of statistical parameter-based TTS systems in our prior works. However, a simple feed-forward deep neural network (DNN) with a f

Artificial IntelligenceComputer Science
3
논문|인용수 38·2019
ExcitNet Vocoder: A Neural Excitation Model for Parametric Speech Synthesis Systems
Eunwoo Song, Kyungguen Byun, Hong-Goo Kang

This paper proposes a WaveNet-based neural excitation model (ExcitNet) for statistical parametric speech synthesis systems. Conventional WaveNet-based neural vocoding systems significantly improve the perceptual quality of synthesized speech by statistically generating a time sequence of speech waveforms through an auto-regressive framework. However, they often suffer from noisy outputs because of the difficulties in capturing the complicated time-varying nature of speech signals. To improve mod

Artificial IntelligenceComputer Science
4
논문|인용수 13·2021
Improved Parallel Wavegan Vocoder with Perceptually Weighted Spectrogram Loss
Eunwoo Song, Ryuichi Yamamoto, Min-Jae Hwang, Jinseob Kim, Ohsung Kwon, Jae-Min Kim

This paper proposes a spectral-domain perceptual weighting technique for Parallel WaveGAN-based text-to-speech (TTS) systems. The recently proposed Parallel WaveGAN vocoder successfully generates waveform sequences using a fast non-autoregressive WaveNet model. By employing multi-resolution short-time Fourier transform (MR-STFT) criteria with a generative adversarial network, the light-weight convolutional networks can be effectively trained without any distillation process. To further improve t

Signal ProcessingComputer Science
5
논문|인용수 9·2015
Improved time-frequency trajectory excitation modeling for a statistical parametric speech synthesis system
Eunwoo Song, Young-Sun Joo, Hong-Goo Kang

This paper proposes an improved time-frequency trajectory excitation (TFTE) modeling method for a statistical parametric speech synthesis system. The proposed approach overcomes the dimensional variation problem of the training process caused by the inherent nature of the pitch-dependent analysis paradigm. By reducing the redundancies of the parameters using predicted average block coefficients (PABC), the proposed algorithm efficiently models excitation, even if its dimension is varied. Objecti

Artificial IntelligenceComputer Science
6
논문|인용수 8·2015
Deep neural network-based statistical parametric speech synthesis system using improved time-frequency trajectory excitation model
Eunwoo Song, Hong-Goo Kang
Artificial IntelligenceComputer Science
7
논문|인용수 7·2020
Neural Text-to-Speech with a Modeling-by-Generation Excitation Vocoder
Eunwoo Song, Min-Jae Hwang, Ryuichi Yamamoto, Jinseob Kim, Ohsung Kwon, Jae-Min Kim

This paper proposes a modeling-by-generation (MbG) excitation vocoder for a neural text-to-speech (TTS) system. Recently proposed neural excitation vocoders can realize qualified waveform generation by combining a vocal tract filter with a WaveNet-based glottal excitation generator. However, when these vocoders are used in a TTS system, the quality of synthesized speech is often degraded owing to a mismatch between training and synthesis steps. Specifically, the vocoder is separately trained fro

Artificial IntelligenceComputer Science
8
논문|인용수 5·2022
TTS-by-TTS 2: Data-Selective Augmentation for Neural Speech Synthesis Using Ranking Support Vector Machine with Variational Autoencoder
Eunwoo Song, Ryuichi Yamamoto, Ohsung Kwon, Chanho Song, Min-Jae Hwang, Suhyeon Oh, Hyun-Wook Yoon, Jinseob Kim, Jae-Min Kim
Interspeech 2022

Recent advances in synthetic speech quality have enabled us to train text-to-speech (TTS) systems by using synthetic corpora. However, merely increasing the amount of synthetic data is not always advantageous for improving training efficiency. Our aim in this study is to selectively choose synthetic data that are beneficial to the training process. In the proposed method, we first adopt a variational autoencoder whose posterior distribution is utilized to extract latent features representing aco

Artificial IntelligenceComputer Science
9
preprint|인용수 4·2024
Improved Parallel Wavegan Vocoder With Perceptually Weighted Spectrogram Loss
Eunwoo Song, Ryuichi Yamamoto, Min-Jae Hwang, Jinseob Kim, Ohsung Kwon, Jae-Min Kim
IEEE RESOURCE CENTERSOA

This paper proposes a spectral-domain perceptual weighting technique for Parallel WaveGAN-based text-to-speech (TTS) systems. The recently proposed Parallel WaveGAN vocoder successfully generates waveform sequences using a fast non-autoregressive WaveNet model. By employing multi-resolution short-time Fourier transform (MR-STFT) criteria with a generative adversarial network, the light-weight convolutional networks can be effectively trained without any distillation process. To further improve t

Signal ProcessingComputer Science
10
논문|인용수 4·2016
Improved Time-Frequency Trajectory Excitation Vocoder for DNN-Based Speech Synthesis
Eunwoo Song, Frank K. Soong, Hong-Goo Kang
Artificial IntelligenceComputer Science
11
논문|인용수 3·2020
Speaker-Adaptive Neural Vocoders for Parametric Speech Synthesis Systems
Eunwoo Song, Jinseob Kim, Kyungguen Byun, Hong-Goo Kang

This paper proposes speaker-adaptive neural vocoders for parametric text-to-speech (TTS) systems. Recently proposed WaveNet-based neural vocoding systems successfully generate a time sequence of speech signal with an autoregressive framework. However, it remains a challenge to synthesize high-quality speech when the amount of a target speaker's training data is insufficient. To generate more natural speech signals with the constraint of limited training data, we propose a speaker adaptation task

Artificial IntelligenceComputer Science
12
preprint|인용수 3·2021
Improved parallel WaveGAN vocoder with perceptually weighted spectrogram loss
Eunwoo Song, Ryuichi Yamamoto, Min-Jae Hwang, Jinseob Kim, Ohsung Kwon, Jae-Min Kim
arXiv (Cornell University)OA

This paper proposes a spectral-domain perceptual weighting technique for Parallel WaveGAN-based text-to-speech (TTS) systems. The recently proposed Parallel WaveGAN vocoder successfully generates waveform sequences using a fast non-autoregressive WaveNet model. By employing multi-resolution short-time Fourier transform (MR-STFT) criteria with a generative adversarial network, the light-weight convolutional networks can be effectively trained without any distillation process. To further improve t

Signal ProcessingComputer Science
13
논문|인용수 2·2017
Perceptual quality and modeling accuracy of excitation parameters in DLSTM-based speech synthesis systems
Eunwoo Song, Frank K. Soong, Hong-Goo Kang
2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)

This paper investigates how the perceptual quality of the synthesized speech is affected by reconstruction errors in excitation signals generated by a deep learning-based statistical model. In this framework, the excitation signal obtained by an LPC inverse filter is first decomposed into harmonic and noise components using an improved time-frequency trajectory excitation (ITFTE) scheme, then they are trained and generated by a deep long short-term memory (DLSTM)-based speech synthesis system. B

Artificial IntelligenceComputer Science
14
preprint|인용수 2·2018
ExcitNet vocoder: A neural excitation model for parametric speech synthesis systems
Eunwoo Song, Kyungguen Byun, Hong-Goo Kang
arXiv (Cornell University)OA

This paper proposes a WaveNet-based neural excitation model (ExcitNet) for statistical parametric speech synthesis systems. Conventional WaveNet-based neural vocoding systems significantly improve the perceptual quality of synthesized speech by statistically generating a time sequence of speech waveforms through an auto-regressive framework. However, they often suffer from noisy outputs because of the difficulties in capturing the complicated time-varying nature of speech signals. To improve mod

Artificial IntelligenceComputer Science
15
논문|인용수 2·2016
Multi-class learning algorithm for deep neural network-based statistical parametric speech synthesis
Eunwoo Song, Hong-Goo Kang

This paper proposes a multi-class learning (MCL) algorithm for a deep neural network (DNN)-based statistical parametric speech synthesis (SPSS) system. Although the DNN-based SPSS system improves the modeling accuracy of statistical parameters, its synthesized speech is often muffled because the training process only considers the global characteristics of the entire set of training data, but does not explicitly consider any local variations. We introduce a DNN-based context clustering algorithm

Artificial IntelligenceComputer Science

대표 연구 분야

Artificial IntelligenceSignal ProcessingMolecular BiologyElectrical and Electronic EngineeringObstetrics and GynecologyBiomedical Engineering

송은우 교수의 연구를 Nubint에서 더 깊이 살펴보세요

이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.