오영환 교수
Young-Hwan Oh
KAIST 전산학부 · 컴퓨터과학
연구실 소개
오영환 교수의 연구실은 음성인식, 감정인식, 음성변환, 소스 분離 등 음성 신호 처리의 핵심 기술을 연구하고 있습니다. 특히 서비스 로봇이 다양한 사용자와 상호작용할 수 있도록 감정 인식 정확도를 높이기 위한 특징 벡터 분류 기법과, 단일 채널 음성에서의 음원 분리 기술에 중점을 두고 있습니다. 또한 음성의 동적 특성과 분포를 효과적으로 모델링하기 위해 은닉 마르코프 모델, 독립성 기반 기저 함수, 스펙트럼 왜곡 모델 등 다양한 수학적 프레임워크를 응용합니다. 연구는 실용적 응용을 고려한 객관적·주관적 평가를 기반으로 한 정밀한 음성 처리 기술 개발을 목표로 합니다.
연구 현황
연구 성과 추이
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
주요 논문
15This paper proposes an efficient feature vector classification for Speech Emotion Recognition (SER) in service robots. Since service robots interact with diverse users who are in various emotional states, two important issues should be addressed: acoustically similar characteristics between emotions and variable speaker characteristics due to different user speaking styles. Each of these issues may cause a substantial amount of overlap between emotion models in feature vector space, thus decreas
We present a new technique for achieving blind source separation when given only a single-channel recording. The main idea is based on exploiting the inherent time structure of sound sources by learning a priori sets of time-domain basis functions that encode the sources in a statistically efficient manner. We derive a learning algorithm using a maximum likelihood approach given the observed single-channel data and sets of basis functions. For each time point, we infer the source parameters and
This paper proposes a new voice conversion technique based on hidden Markov model (HMM) for modeling of speaker’s dynamic characteristics. The basic idea of this technique is to use state transition probability as speaker’s dynamic characteristics and have conversion rule at each state of HMM. A couple of methods is developed for creating state-dependent conversion rule. One uses source speaker’s spectral dynamics and the other uses target speaker’s. The experimental results showed that the prop
We apply independent component analysis for extracting an optimal basis to the problem of finding efficient features for a speaker. The basis functions learned by the algorithm are oriented and localized in both space and frequency, bearing a resemblance to Gabor functions. The speech segments are assumed to be generated by a linear combination of the basis functions, thus the distribution of speech segments of a speaker is modeled by a basis, which is calculated so that each component should be
The performance of a speech recognition system degrades rapidly in the presence of ambient noise. To reduce the degradation, a degradation model is proposed which represents the spectral changes in a speech signal uttered in a noisy environment. The model uses frequency warping and amplitude scaling of each frequency band to simulate the variations of formant location, formant bandwidth, pitch, spectral tilt and energy in each frequency band by the Lombard effect. Another Lombard effect-the vari
In this letter, we propose a new trajectory model for characterizing segmental features and their interaction based upon a general framework of hidden Markov models. Each segment, a sequence of frame vectors, is represented by a trajectory of observed vector sequences. This trajectory replaces the frame features in the segment and becomes the input of the segmental hidden Markov models (HMM's). In our approach, we adopt polynomial trajectory modeling to represent the trajectories using a new des
This paper proposes a new Speech Emotion Recognition (SER) framework. Compared to the speaker-independent emotion models, speaker-adapted models constructed by using a speaker's emotional speech data can represent the speaker's emotional characteristics more precisely, thus improving SER accuracy. However, it is hard to collect a sufficient amount of personal emotional data at once. For this reason, we propose an MLLR-based online speaker adaptation technique using accumulated personal data. Com
Our goal is to extract multiple source signals when only a single observation channel is available. We propose a new signal separation algorithm based on a subspace decomposition. The observation is transformed into subspaces of interest with different sets of basis functions. A flexible model for density estimation allows an accurate modeling of the distributions of the source signals in the subspaces, and we develop a filtering technique using a maximum likelihood (ML) approach to match the ob
This paper proposes an online speaker segmentation approach based on Gaussian Mixture Model (GMM) adaptation for spoken document retrieval. In the conventional approach using the Bayesian Information Criterion (BIC), two single Gaussian models are respectively constructed for two divided speech streams in an analysis window, and the dissimilarity between the two models is estimated according to the BIC principle. This approach has been widely applied to speaker segmentation. However, its perform
Speech Emotion Recognition (SER) is an important area of research in speech processing that aims to identify and classify emotional states conveyed through speech signals. Recent studies have shown considerable performance in SER by exploiting deep contextualized speech representations from self-supervised learning (SSL) models. However, SSL models pre-trained on clean speech data may not perform well on emotional speech data due to the domain shift problem. To address this problem, this paper p
This paper introduces a prosodic phrasing method in Korean to improve the naturalness of speech synthesis, especially in textto-speech conversion. In prosodic phrasing, it is necessary to understand the structure of a sentence through a language processing procedure, such as POS tagging and parsing, since syntactic structure correlates better with the prosodic structure of speech than with other factors. In this paper, the prosodic phrasing procedure is treated from two perspectives: dependency
대표 연구 분야
오영환 교수의 연구를 Nubint에서 더 깊이 살펴보세요
이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.