Skip to main content

Young-Hwan Oh

Korea Advanced Institute of Science and Technology · Computer Science

About the Lab

Professor Young-Hwan Oh's research lab specializes in speech signal processing, with a focus on speech emotion recognition, voice conversion, blind source separation, and robust speech recognition in noisy environments. The lab develops advanced machine learning and statistical models—such as hidden Markov models, independent component analysis, and trajectory-based feature modeling—to address challenges in speaker characterization, acoustic feature representation, and source separation from single-channel recordings. Key research directions include modeling speaker dynamics, enhancing speech recognition robustness under degradation, and designing efficient, statistically optimal basis functions for speech features.

speech emotion recognitionvoice conversionblind source separationrobust speech recognitiontrajectory modeling

Research Overview

Papers
101
Total Citations
544
Papers (5y)
7
Primary Field
Computer Science

Research Output Trend

Figures are computed from collected data and may differ slightly.

Publications per year (5y)
7total
2017
2019
2023
2024
2025
Citations per year (5y)
19total
20172019202320242025

Selected Papers

15
1
Article|100 citations·2009
Feature vector classification based speech emotion recognition for service robots
Jeong‐Sik Park, Ji‐Hwan Kim, Yung‐Hwan Oh
SJR Q1IEEE Transactions on Consumer Electronics

This paper proposes an efficient feature vector classification for Speech Emotion Recognition (SER) in service robots. Since service robots interact with diverse users who are in various emotional states, two important issues should be addressed: acoustically similar characteristics between emotions and variable speaker characteristics due to different user speaking styles. Each of these issues may cause a substantial amount of overlap between emotion models in feature vector space, thus decreas

Experimental and Cognitive PsychologyPsychology
2
Article|57 citations·2003
Single-channel signal separation using time-domain basis functions
Gil‐Jin Jang, Te-Won Lee, Yung‐Hwan Oh
SJR Q1IEEE Signal Processing Letters

We present a new technique for achieving blind source separation when given only a single-channel recording. The main idea is based on exploiting the inherent time structure of sound sources by learning a priori sets of time-domain basis functions that encode the sources in a statistically efficient manner. We derive a learning algorithm using a maximum likelihood approach given the observed single-channel data and sets of basis functions. For each time point, we infer the source parameters and

Signal ProcessingComputer Science
3
Article|57 citations·1999
Tree-based modeling of prosodic phrasing and segmental duration for Korean TTS systems
Sang-Ho Lee, Yung‐Hwan Oh
SJR Q1Speech Communication
Experimental and Cognitive PsychologyPsychology
4
Article|44 citations·1997
Hidden Markov model based voice conversion using dynamic characteristics of speaker
Eun-Kyoung Kim, Sangho Lee, Yung‐Hwan Oh

This paper proposes a new voice conversion technique based on hidden Markov model (HMM) for modeling of speaker’s dynamic characteristics. The basic idea of this technique is to use state transition probability as speaker’s dynamic characteristics and have conversion rule at each state of HMM. A couple of methods is developed for creating state-dependent conversion rule. One uses source speaker’s spectral dynamics and the other uses target speaker’s. The experimental results showed that the prop

Artificial IntelligenceComputer Science
5
Article|41 citations·2002
Learning statistically efficient features for speaker recognition
Gil‐Jin Jang, Te-Won Lee, Yung‐Hwan Oh
SJR Q1Neurocomputing
Signal ProcessingComputer Science
6
Article|19 citations·2002
Learning statistically efficient features for speaker recognition
Gil‐Jin Jang, Te- Won Lee, Yung-Hwan Oh

We apply independent component analysis for extracting an optimal basis to the problem of finding efficient features for a speaker. The basis functions learned by the algorithm are oriented and localized in both space and frequency, bearing a resemblance to Gabor functions. The speech segments are assumed to be generated by a linear combination of the basis functions, thus the distribution of speech segments of a speaker is modeled by a basis, which is calculated so that each component should be

Signal ProcessingComputer Science
7
Article|16 citations·2002
Lombard effect compensation and noise suppression for noisy Lombard speech recognition
Sang-Mun Chi, Yung-Hwan Oh

The performance of a speech recognition system degrades rapidly in the presence of ambient noise. To reduce the degradation, a degradation model is proposed which represents the spectral changes in a speech signal uttered in a noisy environment. The model uses frequency warping and amplitude scaling of each frequency band to simulate the variations of formant location, formant bandwidth, pitch, spectral tilt and energy in each frequency band by the Lombard effect. Another Lombard effect-the vari

Signal ProcessingComputer Science
8
Article|14 citations·2000
A segmental-feature HMM for speech pattern modeling
Youngsun Yun, Yung‐Hwan Oh
SJR Q1IEEE Signal Processing Letters

In this letter, we propose a new trajectory model for characterizing segmental features and their interaction based upon a general framework of hidden Markov models. Each segment, a sequence of frame vectors, is represented by a trajectory of observed vector sequences. This trajectory replaces the frame features in the segment and becomes the input of the segmental hidden Markov models (HMM's). In our approach, we adopt polynomial trajectory modeling to represent the trajectories using a new des

Artificial IntelligenceComputer Science
9
Article|13 citations·2012
Speaker-Characterized Emotion Recognition using Online and Iterative Speaker Adaptation
Jae-Bok Kim, Jeong‐Sik Park, Yung‐Hwan Oh
SJR Q1Cognitive Computation
Experimental and Cognitive PsychologyPsychology
10
Article|13 citations·2011
On-line speaker adaptation based emotion recognition using incremental emotional information
Jae-Bok Kim, Jeong‐Sik Park, Yung‐Hwan Oh

This paper proposes a new Speech Emotion Recognition (SER) framework. Compared to the speaker-independent emotion models, speaker-adapted models constructed by using a speaker's emotional speech data can represent the speaker's emotional characteristics more precisely, thus improving SER accuracy. However, it is hard to collect a sufficient amount of personal emotional data at once. For this reason, we propose an MLLR-based online speaker adaptation technique using accumulated personal data. Com

Experimental and Cognitive PsychologyPsychology
11
Article|12 citations·2004
A subspace approach to single channel signal separation using maximum likelihood weighting filters
Gil‐Jin Jang, Te-Won Lee, Yung‐Hwan Oh
2003 IEEE International Conference on Acoustics, Speech, and Signal Processing, 2003. Proceedings. (ICASSP '03).

Our goal is to extract multiple source signals when only a single observation channel is available. We propose a new signal separation algorithm based on a subspace decomposition. The observation is transformed into subspaces of interest with different sets of basis functions. A flexible model for density estimation allows an accurate modeling of the distributions of the source signals in the subspaces, and we develop a filtering technique using a maximum likelihood (ML) approach to match the ob

Signal ProcessingComputer Science
12
Article|12 citations·2002
A segmental-feature HMM for continuous speech recognition based on a parametric trajectory model
Young-Sun Yun, Yung‐Hwan Oh
SJR Q1Speech Communication
Artificial IntelligenceComputer Science
13
Article|11 citations·2010
GMM adaptation based online speaker segmentation for spoken document retrieval
Kyung-Mi Park, Jeong‐Sik Park, Yung‐Hwan Oh
SJR Q1IEEE Transactions on Consumer Electronics

This paper proposes an online speaker segmentation approach based on Gaussian Mixture Model (GMM) adaptation for spoken document retrieval. In the conventional approach using the Bayesian Information Criterion (BIC), two single Gaussian models are respectively constructed for two divided speech streams in an analysis window, and the dissimilarity between the two models is estimated according to the BIC principle. This approach has been widely applied to speaker segmentation. However, its perform

Artificial IntelligenceComputer Science
14
Article|9 citations·2023
Improving speech emotion recognition by fusing self-supervised learning and spectral features via mixture of experts
Jonghwan Hyeon, Yung‐Hwan Oh, Youngjun Lee, Ho‐Jin Choi
SJR Q2Data & Knowledge EngineeringOA

Speech Emotion Recognition (SER) is an important area of research in speech processing that aims to identify and classify emotional states conveyed through speech signals. Recent studies have shown considerable performance in SER by exploiting deep contextualized speech representations from self-supervised learning (SSL) models. However, SSL models pre-trained on clean speech data may not perform well on emotional speech data due to the domain shift problem. To address this problem, this paper p

Artificial IntelligenceComputer Science
15
Article|8 citations·1999
Prosodic phrasing in korean, determine governor, and then split or not
Yeon-Jun Kim, Heo-Jin Byeon, Yung‐Hwan Oh

This paper introduces a prosodic phrasing method in Korean to improve the naturalness of speech synthesis, especially in textto-speech conversion. In prosodic phrasing, it is necessary to understand the structure of a sentence through a language processing procedure, such as POS tagging and parsing, since syntactic structure correlates better with the prosodic structure of speech than with other factors. In this paper, the prosodic phrasing procedure is treated from two perspectives: dependency

Artificial IntelligenceComputer Science

Research Areas

Artificial IntelligenceSignal ProcessingAerospace EngineeringComputer Vision and Pattern RecognitionExperimental and Cognitive PsychologyControl and Systems Engineering

Dive deeper into Young-Hwan Oh's research on Nubint

Open this lab's papers in the app to read with AI, summarize, and cite in your writing.