Skip to main content

강홍구 교수

Hong-Goo Kang

연세대학교 전기전자공학부 · 컴퓨터과학

연구실 소개

강홍구 교수의 연구실은 음성과 영상 간의 다중모달 통합 인식, 음성 정보 기반의 생체인식 및 병리성 음성 진단 등 음성 및 시각 정보를 융합한 첨단 음성처리 기술을 중심으로 연구를 진행하고 있습니다. 특히, 음성-영상 동기화, 음성 내재화된 워터마킹, 감정 제어 기반 텍스트-음성 합성, 병리성 음성의 특성 분석 등 응용 분야에 특화된 기초 및 응용 연구를 동시에 추구하고 있습니다. 연구는 신경망 기반의 지능형 음성 모델링과 다중모달 학습 기반의 정밀한 음성 분리 기술을 핵심으로 삼고 있습니다.

다중모달 음성처리병리성 음성 진단감정 제어 TTS음성-영상 동기화생체인식 기반 음성 분리

연구 현황

논문 수
249
총 인용 수
1,661
최근 5년 논문
69
주요 분야
컴퓨터과학

연구 성과 추이

표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.

5개년 연도별 논문 게재 수
69총합
2021
2022
2023
2024
2025
5개년 연도별 피인용 수
166총합
20212022202320242025

주요 논문

15
1
논문|인용수 110·2019
Perfect Match: Improved Cross-modal Embeddings for Audio-visual Synchronisation
Soo-Whan Chung, Joon Son Chung, Hong-Goo Kang
OA

This paper proposes a new strategy for learning powerful cross-modal embeddings for audio-to-video synchronisation. Here, we set up the problem as one of cross-modal retrieval, where the objective is to find the most relevant audio segment given a short video clip. The method builds on the recent advances in learning representations from cross-modal self-supervision. The main contributions of this paper are as follows: (1) we propose a new learning strategy where the embeddings are learnt via a

Signal ProcessingComputer Science
2
논문|인용수 92·2017
SVD-Based Adaptive QIM Watermarking on Stereo Audio Signals
Min-Jae Hwang, JeeSok Lee, MiSuk Lee, Hong-Goo Kang
SJR Q1IEEE Transactions on Multimedia

This paper proposes a blind digital audio water- marking algorithm that utilizes the quantization index modulation (QIM) and the singular value decomposition (SVD) of stereo audio signals. Conventional SVD-based blind audio watermarking algorithms lack physical interpretation since the matrix construction method for the input matrix for SVD is heuristically defined. However, in the proposed approach, because the SVD is directly applied to the stereo input signals, the resulting decomposed elemen

Computer Vision and Pattern RecognitionComputer Science
3
논문|인용수 79·2020
Emotional Speech Synthesis with Rich and Granularized Control
Seyun Um, Sangshin Oh, Kyungguen Byun, Inseon Jang, Chunghyun Ahn, Hong-Goo Kang

This paper proposes an effective emotion control method for an end-to-end text-to-speech (TTS) system. To flexibly control the distinct characteristic of a target emotion category, it is essential to determine embedding vectors representing the TTS input. We introduce an inter-to-intra emotional distance ratio algorithm to the embedding vectors that can minimize the distance to the target emotion category while maximizing its distance to the other emotion categories. To further enhance the expre

Artificial IntelligenceComputer Science
4
논문|인용수 73·2017
Effective Spectral and Excitation Modeling Techniques for LSTM-RNN-Based Speech Synthesis Systems
Eunwoo Song, Frank K. Soong, Hong-Goo Kang
SJR Q1IEEE/ACM Transactions on Audio Speech and Language Processing

In this paper, we report research results on modeling the parameters of an improved time-frequency trajectory excitation (ITFTE) and spectral envelopes of an LPC vocoder with a long short-term memory (LSTM)-based recurrent neural network (RNN) for high-quality text-to-speech (TTS) systems. The ITFTE vocoder has been shown to significantly improve the perceptual quality of statistical parameter-based TTS systems in our prior works. However, a simple feed-forward deep neural network (DNN) with a f

Artificial IntelligenceComputer Science
5
논문|인용수 72·2020
FaceFilter: Audio-Visual Speech Separation Using Still Images
Soo-Whan Chung, Soyeon Choe, Joon Son Chung, Hong-Goo Kang
OA

The objective of this paper is to separate a target speaker's speech from a mixture of two speakers using a deep audio-visual speech separation network. Unlike previous works that used lip movement on video clips or pre-enrolled speaker information as an auxiliary conditional feature, we use a single face image of the target speaker. In this task, the conditional feature is obtained from facial appearance in cross-modal biometric task, where audio and visual identity representations are shared i

Signal ProcessingComputer Science
6
논문|인용수 47·2013
An Investigation of Vocal Tract Characteristics for Acoustic Discrimination of Pathological Voices
Jung-Won Lee, Hong-Goo Kang, Jeung-Yoon Choi, Young-Ik Son
SJR Q2BioMed Research InternationalOA

This paper investigates the effectiveness of measures related to vocal tract characteristics in classifying normal and pathological speech. Unlike conventional approaches that mainly focus on features related to the vocal source, vocal tract characteristics are examined to determine if interaction effects between vocal folds and the vocal tract can be used to detect pathological speech. Especially, this paper examines features related to formant frequencies to see if vocal tract characteristics

PhysiologyMedicine
7
논문|인용수 46·2018
A Deep Learning-based Stress Detection Algorithm with Speech Signal
Hye-Won Han, Kyunggeun Byun, Hong-Goo Kang

In this paper, we propose a deep learning-based psychological stress detection algorithm using speech signals. With increasing demands for communication between human and intelligent systems, automatic stress detection is becoming an interesting research topic. Stress can be reliably detected by measuring the level of specific hormones (e.g., cortisol), but this is not a convenient method for the detection of stress in human-machine interactions. The proposed algorithm first extracts mel-filterb

Experimental and Cognitive PsychologyPsychology
8
논문|인용수 44·2019
An Effective Style Token Weight Control Technique for End-to-End Emotional Speech Synthesis
Ohsung Kwon, Inseon Jang, Chunghyun Ahn, Hong-Goo Kang
SJR Q1IEEE Signal Processing Letters

In this letter, we propose a high-quality emotional speech synthesis system, using emotional vector space, i.e., the weighted sum of global style tokens (GSTs). Our previous research verified the feasibility of GST-based emotional speech synthesis in an end-to-end text-to-speech synthesis framework. However, selecting appropriate reference audio (RA) signals to extract emotion embedding vectors to the specific types of target emotions remains problematic. To ameliorate the selection problem, we

Artificial IntelligenceComputer Science
9
논문|인용수 41·2018
Phase-Sensitive Joint Learning Algorithms for Deep Learning-Based Speech Enhancement
Jinkyu Lee, Jan Skoglund, Turaj Zakizadeh Shabestary, Hong-Goo Kang
SJR Q1IEEE Signal Processing Letters

This letter presents a phase-sensitive joint learning algorithm for single-channel speech enhancement. Although a deep learning framework that estimates the time-frequency (T-F) domain ideal ratio masks demonstrates a strong performance, it is limited in the sense that the enhancement process is performed only in the magnitude domain, while the phase spectra remain unchanged. Thus, recent studies have been conducted to involve phase spectra in speech enhancement systems. A phase-sensitive mask (

Signal ProcessingComputer Science
10
논문|인용수 40·2020
Seeing Voices and Hearing Voices: Learning Discriminative Embeddings Using Cross-Modal Self-Supervision
Soo-Whan Chung, Hong-Goo Kang, Joon Son Chung
OA

The goal of this work is to train discriminative cross-modal embeddings without access to manually annotated data. Recent advances in self-supervised learning have shown that effective representations can be learnt from natural cross-modal synchrony. We build on earlier work to train embeddings that are more discriminative for uni-modal downstream tasks. To this end, we propose a novel training strategy that not only optimises metrics across modalities, but also enforces intra-class feature sepa

Signal ProcessingComputer Science
11
논문|인용수 38·2019
ExcitNet Vocoder: A Neural Excitation Model for Parametric Speech Synthesis Systems
Eunwoo Song, Kyungguen Byun, Hong-Goo Kang

This paper proposes a WaveNet-based neural excitation model (ExcitNet) for statistical parametric speech synthesis systems. Conventional WaveNet-based neural vocoding systems significantly improve the perceptual quality of synthesized speech by statistically generating a time sequence of speech waveforms through an auto-regressive framework. However, they often suffer from noisy outputs because of the difficulties in capturing the complicated time-varying nature of speech signals. To improve mod

Artificial IntelligenceComputer Science
12
논문|인용수 19·2003
Improving the transcoding capability of speech coders
Hong-Goo Kang, Hong Kook Kim, R.V. Cox
SJR Q1IEEE Transactions on Multimedia

With the trend of merging various communication networks, a need arises to provide transcoding between different speech coding formats. Presently this means a cross tandem between the two coders in each case. This results in both quality loss and extra delay. A possible alternative is using a bitstream mapping approach that directly converts parameter values. For several standard coders having a similar coding structure, it should be possible to generate comparable or better quality without addi

Computer Vision and Pattern RecognitionComputer Science
13
논문|인용수 16·2016
On pre-filtering strategies for the GCC-PHAT algorithm
Hong-Goo Kang, Michael Graczyk, Jan Skoglund

In this paper, we investigate the impact of the pre-filtering method to generalized cross-correlation (GCC) based direction of arrival (DOA) estimation. The role of pre-filtering is either to emphasize or deemphasize certain frequency components before computing cross power spectrum. However, its impact or relation to environmental variation, e.g., in noisy environments, has not been clearly studied yet. An efficient pre-filter should consider the relative importance of individual frequency comp

Signal ProcessingComputer Science
14
논문|인용수 14·2002
Improving transcoding capability of speech coders in clean and frame erasured channel environments
Hong-Goo Kang, Hong Kook Kim, R.V. Cox

With the trend of merging various networks, a need arises to provide transcoding between different speech coding formats. Presently this means cross tandeming the two coders, but it results in both quality loss and extra delay. A possible alternative is using a bit-stream mapping approach that directly converts parameter values. This paper proposes a bit-stream mapping method between ITU-T Recommendation G.729 and TIA IS-641. Informal listening tests and PSQM scores show that the proposed method

Computer Vision and Pattern RecognitionComputer Science
15
논문|인용수 11·2019
Dry Electrode-Based Body Fat Estimation System with Anthropometric Data for Use in a Wearable Device
Seung-Chul Shin, Jinkyu Lee, Soyeon Choe, Hyuk In Yang, Jihee Min, Ki-Yong Ahn, Justin Y. Jeon, Hong-Goo Kang
SJR Q1SensorsOA

The bioelectrical impedance analysis (BIA) method is widely used to predict percent body fat (PBF). However, it requires four to eight electrodes, and it takes a few minutes to accurately obtain the measurement results. In this study, we propose a faster and more accurate method that utilizes a small dry electrode-based wearable device, which predicts whole-body impedance using only upper-body impedance values. Such a small electrode-based device typically needs a long measurement time due to in

PhysiologyMedicine

대표 연구 분야

Signal ProcessingArtificial IntelligenceComputer Vision and Pattern RecognitionAerospace EngineeringCognitive NeuroscienceEconomics and Econometrics

강홍구 교수의 연구를 Nubint에서 더 깊이 살펴보세요

이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.