Namsoo Kim
Seoul National University · 情報科学
研究室紹介
Professor Namsoo Kim's research lab specializes in speech signal processing and robust speech recognition, with a strong focus on developing advanced techniques for speech enhancement and noise robustness in real-world environments. The lab explores statistical modeling, adaptive filtering, and machine learning approaches—ranging from classical methods like the interacting multiple model (IMM) and sequential EM algorithms to modern deep learning frameworks such as GANs and multi-scale architectures. A key research direction involves leveraging temporal dynamics and multi-resolution representations to improve speech quality and recognition accuracy under nonstationary and adverse acoustic conditions.
Research Overview
Research Output Trend
Figures are computed from collected data and may differ slightly.
Selected Papers
15In this letter, we propose a novel speech enhancement technique based on global soft decision. The proposed approach provides a unified framework for such procedures as speech absence probability (SAP) computation, spectral gain modification, and noise spectrum estimation using the same statistical model assumption. Performances of the proposed enhancement algorithm are evaluated by subjective tests under various environments and show better results compared with the IS-127 standard enhancement
Sequential approaches are proposed to compensate for the effects of the nonstationary environment for robust speech recognition. Unlike the batch approaches, the proposed methods derive a different parameter estimate for each time using the sequential expectation maximization (EM) algorithm. Moreover, we also propose the forward-backward estimation scheme as an improvement of the sequential parameter estimation.
We propose a new approach to environmental parameter estimation for robust speech recognition in adverse conditions. The proposed method is based on the interacting multiple model (IMM) technique widely used in the area of multiple target tracking. Through a number of continuous digit recognition experiments, we can find the effectiveness of the IMM-based approach in slowly evolving environment conditions.
Recently, generative adversarial networks (GANs) have been successfully applied to speech enhancement. However, there still remain two issues that need to be addressed: (1) GAN-based training is typically unstable due to its non-convex property, and (2) most of the conventional methods do not fully take advantage of the speech characteristics, which could result in a sub-optimal solution. In order to deal with these problems, we propose a progressive generator that can handle the speech in a mul
문화역사활동이론(CHAT)은 장기간에 걸쳐서 일어난 일련의 목표 지향적이며 인공물이 매개하는 집단적 행위로 이루어진 활동을 기본적인 분석 단위로 삼는다. 하나의 활동 체계는 주체, 목표, 인공물, 규칙, 공동체 그리고 분업 등 여섯 가지의 기본 요소로 구성되어 있으며 활동이 일어나는 과정에서 발생하는 딜레마들은 개별 요소 내부의, 요소 간의 혹은 각기 다른 활동 체계 간의 모순으로 설명하고 이 모순들은 변화와 개선의 실마리를 찾는 지점으로 본다. 본 논문은 문화역사활동이론이 제안하는 활동 체계를 틀로 삼아 서울형 혁신학교 사업에 참여한 지 1년차인 A 중학교의 수업 혁신 활동을 기술하고 분석하였다. 한 학기 동안 진행된 참여 관찰을 통하여 얻은 자료들을 수업 혁신 활동 체계를 시작하게 된 매개 개념과 목표, 수업 혁신 활동의 주체와 인공물, 공동체와 규칙 그리고 분업(과 협업) 등의 요소로 구분하였다. 수업에서 발견된 문제적 상황이자 개선을 위해 해법이 필요한 지점은 기존의 수업 활
In this letter, we propose a novel approach to feature compensation for robust speech recognition in noisy environments. We employ the switching linear dynamic model (SLDM) as a parametric model for the clean speech distribution, which enables us to exploit temporal correlations inherent in speech signals. Both the background noise and clean speech components are simultaneously estimated by means of the interacting multiple model (IMM) algorithm.
및 시사점본 연구는 국외 생태학교 관련 사업이 국내 환경보전시범학교 실시에 주는 시사점을 얻기 위해 실시되었으며, 운영 목적, 선발 방식, 운영 방식, 지원 방식, 평가 방식, 주요 장애 등에 주안점을 두어 국가별 특징과 국내 사례와의 비교 분석을 수행하여 시사점을 얻고자 하였다. 외국의 경우, 지속가능발전교육과 환경교육의 통합을 꾀하거나 또는 환경교육의 방향을 지속가능발전교육의 지향을 담아 재정향하는 노력이 진행되고 있다. 이들은 국가 수준의 환경교육 전략에 이러한 내용을 담거나 또는 기존의 녹색학교, 생태 학교 등의 이름으로 지칭되던 학교 수준의 지원 사업을 지속가능한 학교 등의 이름으로 개칭하고 학교 내 환경교육 및 지속가능발전교육의 중요성을 함께 강조하고 있다. 학생 교육과 교사에 대한 신뢰 회복을 위해 학교와 학부모의 협력 관계가 요구되며 학생이 지역 사회 구성원으로서의 자아를 확립하기 위해 학교와 지역사회의 파트너쉽 형성이 중요한데(Tett, 2004) 학교 전체 접근은
One of the most popular approaches to parameter adaptation in hidden Markov model (HMM) based systems is the maximum likelihood linear regression (MLLR) technique. In this letter, we extend MLLR to factored MLLR (FMLLR) in which the MLLR parameters depend on a continuous-valued control vector. Since it is practically impossible to estimate the MLLR parameters for each control vector separately, we propose a compact parametric form of the MLLR parameters. In the proposed approach, each MLLR param
본 연구는 문화역사활동이론(Cultural Historical Activity Theory: CHAT)의 활동 체계를 분석틀로 삼아 수업 전문성 신장 활동인 수업 장학, 수업 컨설팅, 수업 비평을 살펴보았다. 그 과정에서 각 활동을 구분하는 기준 질문을 도출하였으며 그 질문들을 바탕으로 세 활동의 관계를 통시적인 관점과 공시적인 관점에서 설명하고자 했다. CHAT은 수업 전문성 신장의 단위를 교사 개인이 아니라 수업을 중심으로 대화를 나누는 집단으로 확장하여 보기를 제안한다. 이 관점에서 본다면 각 활동이 이루어지는 공동체 내에서 역할 바꿈 가능성, 인공물과 규칙의 합의 가능성은 수업 전문성 신장 활동의 성격과 의미를 파악하는 데 중요한 기준 질문이 된다. 이러한 기준 질문을 토대로 세 활동의 관계를 통시적으로 보았을 때 각 활동의 등장은 기존의 활동 체계 속의 문제점들을 해결하려는 시도로 볼 수 있으며 우리나라 수업 전문성 신장 활동 체계를 구성하는 각 요소들은 수정되거나 확장되었
We present a novel method to incorporate temporal correlations into a speech recognition system based on conventional hidden Markov models (HMMs). The temporal correlations are considered to be useful for recognition because of the fact that the speech features of the present frame are highly informative about the feature characteristics of neighboring frames. In this paper, by treating these correlations in the form of conditional probability distributions (PDs), we propose a new technique for
Abstract Acoustic data transmission (ADT) forms a branch of the audio data hiding techniques with its capability of communicating data in short-range aerial space between a loudspeaker and a microphone. In this paper, we propose an acoustic data transmission system extending our previous studies and give an in-depth analysis of its performance. The proposed technique utilizes the phases of modulated complex lapped transform (MCLT) coefficients of the audio signal. To achieve a good trade-off bet
We present various methods for estimating a robust output probability distribution (PD) in speech recognition based on the discrete hidden Markov model (HMM). In speech recognition, we encounter the problem of an insufficient amount of training data, which may cause inaccurate modeling of the HMM parameters, especially the output PD's. In this paper, to enhance the robustness of the output PD's with respect to unseen data, we study two approaches: smoothing and tying of the PD's. We introduce a
Recently, the increasing demand for voice-based authentication systems has encouraged researchers to investigate methods for verifying users with short randomized pass-phrases with constrained vocabulary. The conventional i-vector framework, which has been proven to be a state-of-the-art utterance-level feature extraction technique for speaker verification, is not considered to be an optimal method for this task since it is known to suffer from severe performance degradation when dealing with sh