Skip to main content

Hong-Goo Kang

Yonsei University · Computer Science

About the Lab

Professor Hong-Goo Kang's research lab specializes in audio-visual speech processing, speech synthesis, and biometrics, with a strong focus on leveraging deep learning and self-supervised representation learning for cross-modal understanding. Key research directions include audio-to-video synchronization, blind audio watermarking, emotion-controlled text-to-speech systems, and speaker separation using visual cues. The lab also investigates vocal tract characteristics for pathological speech detection, emphasizing the interplay between vocal source and articulatory features.

audio-visual learningtext-to-speech synthesisspeech separationspeech emotion recognitionbiometric speech processing

Research Overview

Papers
249
Total Citations
1,661
Papers (5y)
69
Primary Field
Computer Science

Research Output Trend

Figures are computed from collected data and may differ slightly.

Publications per year (5y)
69total
2021
2022
2023
2024
2025
Citations per year (5y)
166total
20212022202320242025

Selected Papers

15
1
Article|110 citations·2019
Perfect Match: Improved Cross-modal Embeddings for Audio-visual Synchronisation
Soo-Whan Chung, Joon Son Chung, Hong-Goo Kang
OA

This paper proposes a new strategy for learning powerful cross-modal embeddings for audio-to-video synchronisation. Here, we set up the problem as one of cross-modal retrieval, where the objective is to find the most relevant audio segment given a short video clip. The method builds on the recent advances in learning representations from cross-modal self-supervision. The main contributions of this paper are as follows: (1) we propose a new learning strategy where the embeddings are learnt via a

Signal ProcessingComputer Science
2
Article|92 citations·2017
SVD-Based Adaptive QIM Watermarking on Stereo Audio Signals
Min-Jae Hwang, JeeSok Lee, MiSuk Lee, Hong-Goo Kang
SJR Q1IEEE Transactions on Multimedia

This paper proposes a blind digital audio water- marking algorithm that utilizes the quantization index modulation (QIM) and the singular value decomposition (SVD) of stereo audio signals. Conventional SVD-based blind audio watermarking algorithms lack physical interpretation since the matrix construction method for the input matrix for SVD is heuristically defined. However, in the proposed approach, because the SVD is directly applied to the stereo input signals, the resulting decomposed elemen

Computer Vision and Pattern RecognitionComputer Science
3
Article|79 citations·2020
Emotional Speech Synthesis with Rich and Granularized Control
Seyun Um, Sangshin Oh, Kyungguen Byun, Inseon Jang, Chunghyun Ahn, Hong-Goo Kang

This paper proposes an effective emotion control method for an end-to-end text-to-speech (TTS) system. To flexibly control the distinct characteristic of a target emotion category, it is essential to determine embedding vectors representing the TTS input. We introduce an inter-to-intra emotional distance ratio algorithm to the embedding vectors that can minimize the distance to the target emotion category while maximizing its distance to the other emotion categories. To further enhance the expre

Artificial IntelligenceComputer Science
4
Article|73 citations·2017
Effective Spectral and Excitation Modeling Techniques for LSTM-RNN-Based Speech Synthesis Systems
Eunwoo Song, Frank K. Soong, Hong-Goo Kang
SJR Q1IEEE/ACM Transactions on Audio Speech and Language Processing

In this paper, we report research results on modeling the parameters of an improved time-frequency trajectory excitation (ITFTE) and spectral envelopes of an LPC vocoder with a long short-term memory (LSTM)-based recurrent neural network (RNN) for high-quality text-to-speech (TTS) systems. The ITFTE vocoder has been shown to significantly improve the perceptual quality of statistical parameter-based TTS systems in our prior works. However, a simple feed-forward deep neural network (DNN) with a f

Artificial IntelligenceComputer Science
5
Article|72 citations·2020
FaceFilter: Audio-Visual Speech Separation Using Still Images
Soo-Whan Chung, Soyeon Choe, Joon Son Chung, Hong-Goo Kang
OA

The objective of this paper is to separate a target speaker's speech from a mixture of two speakers using a deep audio-visual speech separation network. Unlike previous works that used lip movement on video clips or pre-enrolled speaker information as an auxiliary conditional feature, we use a single face image of the target speaker. In this task, the conditional feature is obtained from facial appearance in cross-modal biometric task, where audio and visual identity representations are shared i

Signal ProcessingComputer Science
6
Article|47 citations·2013
An Investigation of Vocal Tract Characteristics for Acoustic Discrimination of Pathological Voices
Jung-Won Lee, Hong-Goo Kang, Jeung-Yoon Choi, Young-Ik Son
SJR Q2BioMed Research InternationalOA

This paper investigates the effectiveness of measures related to vocal tract characteristics in classifying normal and pathological speech. Unlike conventional approaches that mainly focus on features related to the vocal source, vocal tract characteristics are examined to determine if interaction effects between vocal folds and the vocal tract can be used to detect pathological speech. Especially, this paper examines features related to formant frequencies to see if vocal tract characteristics

PhysiologyMedicine
7
Article|46 citations·2018
A Deep Learning-based Stress Detection Algorithm with Speech Signal
Hye-Won Han, Kyunggeun Byun, Hong-Goo Kang

In this paper, we propose a deep learning-based psychological stress detection algorithm using speech signals. With increasing demands for communication between human and intelligent systems, automatic stress detection is becoming an interesting research topic. Stress can be reliably detected by measuring the level of specific hormones (e.g., cortisol), but this is not a convenient method for the detection of stress in human-machine interactions. The proposed algorithm first extracts mel-filterb

Experimental and Cognitive PsychologyPsychology
8
Article|44 citations·2019
An Effective Style Token Weight Control Technique for End-to-End Emotional Speech Synthesis
Ohsung Kwon, Inseon Jang, Chunghyun Ahn, Hong-Goo Kang
SJR Q1IEEE Signal Processing Letters

In this letter, we propose a high-quality emotional speech synthesis system, using emotional vector space, i.e., the weighted sum of global style tokens (GSTs). Our previous research verified the feasibility of GST-based emotional speech synthesis in an end-to-end text-to-speech synthesis framework. However, selecting appropriate reference audio (RA) signals to extract emotion embedding vectors to the specific types of target emotions remains problematic. To ameliorate the selection problem, we

Artificial IntelligenceComputer Science
9
Article|41 citations·2018
Phase-Sensitive Joint Learning Algorithms for Deep Learning-Based Speech Enhancement
Jinkyu Lee, Jan Skoglund, Turaj Zakizadeh Shabestary, Hong-Goo Kang
SJR Q1IEEE Signal Processing Letters

This letter presents a phase-sensitive joint learning algorithm for single-channel speech enhancement. Although a deep learning framework that estimates the time-frequency (T-F) domain ideal ratio masks demonstrates a strong performance, it is limited in the sense that the enhancement process is performed only in the magnitude domain, while the phase spectra remain unchanged. Thus, recent studies have been conducted to involve phase spectra in speech enhancement systems. A phase-sensitive mask (

Signal ProcessingComputer Science
10
Article|40 citations·2020
Seeing Voices and Hearing Voices: Learning Discriminative Embeddings Using Cross-Modal Self-Supervision
Soo-Whan Chung, Hong-Goo Kang, Joon Son Chung
OA

The goal of this work is to train discriminative cross-modal embeddings without access to manually annotated data. Recent advances in self-supervised learning have shown that effective representations can be learnt from natural cross-modal synchrony. We build on earlier work to train embeddings that are more discriminative for uni-modal downstream tasks. To this end, we propose a novel training strategy that not only optimises metrics across modalities, but also enforces intra-class feature sepa

Signal ProcessingComputer Science
11
Article|38 citations·2019
ExcitNet Vocoder: A Neural Excitation Model for Parametric Speech Synthesis Systems
Eunwoo Song, Kyungguen Byun, Hong-Goo Kang

This paper proposes a WaveNet-based neural excitation model (ExcitNet) for statistical parametric speech synthesis systems. Conventional WaveNet-based neural vocoding systems significantly improve the perceptual quality of synthesized speech by statistically generating a time sequence of speech waveforms through an auto-regressive framework. However, they often suffer from noisy outputs because of the difficulties in capturing the complicated time-varying nature of speech signals. To improve mod

Artificial IntelligenceComputer Science
12
Article|19 citations·2003
Improving the transcoding capability of speech coders
Hong-Goo Kang, Hong Kook Kim, R.V. Cox
SJR Q1IEEE Transactions on Multimedia

With the trend of merging various communication networks, a need arises to provide transcoding between different speech coding formats. Presently this means a cross tandem between the two coders in each case. This results in both quality loss and extra delay. A possible alternative is using a bitstream mapping approach that directly converts parameter values. For several standard coders having a similar coding structure, it should be possible to generate comparable or better quality without addi

Computer Vision and Pattern RecognitionComputer Science
13
Article|16 citations·2016
On pre-filtering strategies for the GCC-PHAT algorithm
Hong-Goo Kang, Michael Graczyk, Jan Skoglund

In this paper, we investigate the impact of the pre-filtering method to generalized cross-correlation (GCC) based direction of arrival (DOA) estimation. The role of pre-filtering is either to emphasize or deemphasize certain frequency components before computing cross power spectrum. However, its impact or relation to environmental variation, e.g., in noisy environments, has not been clearly studied yet. An efficient pre-filter should consider the relative importance of individual frequency comp

Signal ProcessingComputer Science
14
Article|14 citations·2002
Improving transcoding capability of speech coders in clean and frame erasured channel environments
Hong-Goo Kang, Hong Kook Kim, R.V. Cox

With the trend of merging various networks, a need arises to provide transcoding between different speech coding formats. Presently this means cross tandeming the two coders, but it results in both quality loss and extra delay. A possible alternative is using a bit-stream mapping approach that directly converts parameter values. This paper proposes a bit-stream mapping method between ITU-T Recommendation G.729 and TIA IS-641. Informal listening tests and PSQM scores show that the proposed method

Computer Vision and Pattern RecognitionComputer Science
15
Article|11 citations·2019
Dry Electrode-Based Body Fat Estimation System with Anthropometric Data for Use in a Wearable Device
Seung-Chul Shin, Jinkyu Lee, Soyeon Choe, Hyuk In Yang, Jihee Min, Ki-Yong Ahn, Justin Y. Jeon, Hong-Goo Kang
SJR Q1SensorsOA

The bioelectrical impedance analysis (BIA) method is widely used to predict percent body fat (PBF). However, it requires four to eight electrodes, and it takes a few minutes to accurately obtain the measurement results. In this study, we propose a faster and more accurate method that utilizes a small dry electrode-based wearable device, which predicts whole-body impedance using only upper-body impedance values. Such a small electrode-based device typically needs a long measurement time due to in

PhysiologyMedicine

Research Areas

Signal ProcessingArtificial IntelligenceComputer Vision and Pattern RecognitionAerospace EngineeringCognitive NeuroscienceEconomics and Econometrics

Dive deeper into Hong-Goo Kang's research on Nubint

Open this lab's papers in the app to read with AI, summarize, and cite in your writing.