Young-Hwan Oh
Korea Advanced Institute of Science and Technology · 情報科学
研究室紹介
Professor Young-Hwan Oh's research lab specializes in speech signal processing, with a focus on speech emotion recognition, voice conversion, blind source separation, and robust speech recognition in noisy environments. The lab develops advanced machine learning and statistical models—such as hidden Markov models, independent component analysis, and trajectory-based feature modeling—to address challenges in speaker characterization, acoustic feature representation, and source separation from single-channel recordings. Key research directions include modeling speaker dynamics, enhancing speech recognition robustness under degradation, and designing efficient, statistically optimal basis functions for speech features.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15This paper proposes an efficient feature vector classification for Speech Emotion Recognition (SER) in service robots. Since service robots interact with diverse users who are in various emotional states, two important issues should be addressed: acoustically similar characteristics between emotions and variable speaker characteristics due to different user speaking styles. Each of these issues may cause a substantial amount of overlap between emotion models in feature vector space, thus decreas
We present a new technique for achieving blind source separation when given only a single-channel recording. The main idea is based on exploiting the inherent time structure of sound sources by learning a priori sets of time-domain basis functions that encode the sources in a statistically efficient manner. We derive a learning algorithm using a maximum likelihood approach given the observed single-channel data and sets of basis functions. For each time point, we infer the source parameters and
This paper proposes a new voice conversion technique based on hidden Markov model (HMM) for modeling of speaker’s dynamic characteristics. The basic idea of this technique is to use state transition probability as speaker’s dynamic characteristics and have conversion rule at each state of HMM. A couple of methods is developed for creating state-dependent conversion rule. One uses source speaker’s spectral dynamics and the other uses target speaker’s. The experimental results showed that the prop
We apply independent component analysis for extracting an optimal basis to the problem of finding efficient features for a speaker. The basis functions learned by the algorithm are oriented and localized in both space and frequency, bearing a resemblance to Gabor functions. The speech segments are assumed to be generated by a linear combination of the basis functions, thus the distribution of speech segments of a speaker is modeled by a basis, which is calculated so that each component should be
The performance of a speech recognition system degrades rapidly in the presence of ambient noise. To reduce the degradation, a degradation model is proposed which represents the spectral changes in a speech signal uttered in a noisy environment. The model uses frequency warping and amplitude scaling of each frequency band to simulate the variations of formant location, formant bandwidth, pitch, spectral tilt and energy in each frequency band by the Lombard effect. Another Lombard effect-the vari
In this letter, we propose a new trajectory model for characterizing segmental features and their interaction based upon a general framework of hidden Markov models. Each segment, a sequence of frame vectors, is represented by a trajectory of observed vector sequences. This trajectory replaces the frame features in the segment and becomes the input of the segmental hidden Markov models (HMM's). In our approach, we adopt polynomial trajectory modeling to represent the trajectories using a new des
This paper proposes a new Speech Emotion Recognition (SER) framework. Compared to the speaker-independent emotion models, speaker-adapted models constructed by using a speaker's emotional speech data can represent the speaker's emotional characteristics more precisely, thus improving SER accuracy. However, it is hard to collect a sufficient amount of personal emotional data at once. For this reason, we propose an MLLR-based online speaker adaptation technique using accumulated personal data. Com
Our goal is to extract multiple source signals when only a single observation channel is available. We propose a new signal separation algorithm based on a subspace decomposition. The observation is transformed into subspaces of interest with different sets of basis functions. A flexible model for density estimation allows an accurate modeling of the distributions of the source signals in the subspaces, and we develop a filtering technique using a maximum likelihood (ML) approach to match the ob
This paper proposes an online speaker segmentation approach based on Gaussian Mixture Model (GMM) adaptation for spoken document retrieval. In the conventional approach using the Bayesian Information Criterion (BIC), two single Gaussian models are respectively constructed for two divided speech streams in an analysis window, and the dissimilarity between the two models is estimated according to the BIC principle. This approach has been widely applied to speaker segmentation. However, its perform
Speech Emotion Recognition (SER) is an important area of research in speech processing that aims to identify and classify emotional states conveyed through speech signals. Recent studies have shown considerable performance in SER by exploiting deep contextualized speech representations from self-supervised learning (SSL) models. However, SSL models pre-trained on clean speech data may not perform well on emotional speech data due to the domain shift problem. To address this problem, this paper p
This paper introduces a prosodic phrasing method in Korean to improve the naturalness of speech synthesis, especially in textto-speech conversion. In prosodic phrasing, it is necessary to understand the structure of a sentence through a language processing procedure, such as POS tagging and parsing, since syntactic structure correlates better with the prosodic structure of speech than with other factors. In this paper, the prosodic phrasing procedure is treated from two perspectives: dependency