김남수 교수
Namsoo Kim
서울대학교 · 컴퓨터과학
연구실 소개
김남수 교수의 연구실은 음성 인식 및 음성 향상 기술 분야에서 핵심적인 연구를 수행하고 있습니다. 비정상적인 환경에서도 안정적으로 작동하는 음성 처리 기법 개발에 초점을 맞추며, 특히 비모수적 모델과 통계적 추정 기반의 신호 처리 기법을 활용한 혁신적인 알고리즘 설계에 기여하고 있습니다. 최근에는 딥러닝 기반 음성 향상 기술의 안정성과 성능 향상을 위해 다중 해상도 및 다중 척도 기반의 GAN 아키텍처 개발에도 주력하고 있습니다. 또한 교육 현장의 혁신 활동을 분석하는 데 있어 문화역사활동이론을 응용한 교육 기술 연구도 함께 진행하고 있습니다.
연구 현황
연구 성과 추이
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
주요 논문
15In this letter, we propose a novel speech enhancement technique based on global soft decision. The proposed approach provides a unified framework for such procedures as speech absence probability (SAP) computation, spectral gain modification, and noise spectrum estimation using the same statistical model assumption. Performances of the proposed enhancement algorithm are evaluated by subjective tests under various environments and show better results compared with the IS-127 standard enhancement
Sequential approaches are proposed to compensate for the effects of the nonstationary environment for robust speech recognition. Unlike the batch approaches, the proposed methods derive a different parameter estimate for each time using the sequential expectation maximization (EM) algorithm. Moreover, we also propose the forward-backward estimation scheme as an improvement of the sequential parameter estimation.
We propose a new approach to environmental parameter estimation for robust speech recognition in adverse conditions. The proposed method is based on the interacting multiple model (IMM) technique widely used in the area of multiple target tracking. Through a number of continuous digit recognition experiments, we can find the effectiveness of the IMM-based approach in slowly evolving environment conditions.
Recently, generative adversarial networks (GANs) have been successfully applied to speech enhancement. However, there still remain two issues that need to be addressed: (1) GAN-based training is typically unstable due to its non-convex property, and (2) most of the conventional methods do not fully take advantage of the speech characteristics, which could result in a sub-optimal solution. In order to deal with these problems, we propose a progressive generator that can handle the speech in a mul
문화역사활동이론(CHAT)은 장기간에 걸쳐서 일어난 일련의 목표 지향적이며 인공물이 매개하는 집단적 행위로 이루어진 활동을 기본적인 분석 단위로 삼는다. 하나의 활동 체계는 주체, 목표, 인공물, 규칙, 공동체 그리고 분업 등 여섯 가지의 기본 요소로 구성되어 있으며 활동이 일어나는 과정에서 발생하는 딜레마들은 개별 요소 내부의, 요소 간의 혹은 각기 다른 활동 체계 간의 모순으로 설명하고 이 모순들은 변화와 개선의 실마리를 찾는 지점으로 본다. 본 논문은 문화역사활동이론이 제안하는 활동 체계를 틀로 삼아 서울형 혁신학교 사업에 참여한 지 1년차인 A 중학교의 수업 혁신 활동을 기술하고 분석하였다. 한 학기 동안 진행된 참여 관찰을 통하여 얻은 자료들을 수업 혁신 활동 체계를 시작하게 된 매개 개념과 목표, 수업 혁신 활동의 주체와 인공물, 공동체와 규칙 그리고 분업(과 협업) 등의 요소로 구분하였다. 수업에서 발견된 문제적 상황이자 개선을 위해 해법이 필요한 지점은 기존의 수업 활
In this letter, we propose a novel approach to feature compensation for robust speech recognition in noisy environments. We employ the switching linear dynamic model (SLDM) as a parametric model for the clean speech distribution, which enables us to exploit temporal correlations inherent in speech signals. Both the background noise and clean speech components are simultaneously estimated by means of the interacting multiple model (IMM) algorithm.
및 시사점본 연구는 국외 생태학교 관련 사업이 국내 환경보전시범학교 실시에 주는 시사점을 얻기 위해 실시되었으며, 운영 목적, 선발 방식, 운영 방식, 지원 방식, 평가 방식, 주요 장애 등에 주안점을 두어 국가별 특징과 국내 사례와의 비교 분석을 수행하여 시사점을 얻고자 하였다. 외국의 경우, 지속가능발전교육과 환경교육의 통합을 꾀하거나 또는 환경교육의 방향을 지속가능발전교육의 지향을 담아 재정향하는 노력이 진행되고 있다. 이들은 국가 수준의 환경교육 전략에 이러한 내용을 담거나 또는 기존의 녹색학교, 생태 학교 등의 이름으로 지칭되던 학교 수준의 지원 사업을 지속가능한 학교 등의 이름으로 개칭하고 학교 내 환경교육 및 지속가능발전교육의 중요성을 함께 강조하고 있다. 학생 교육과 교사에 대한 신뢰 회복을 위해 학교와 학부모의 협력 관계가 요구되며 학생이 지역 사회 구성원으로서의 자아를 확립하기 위해 학교와 지역사회의 파트너쉽 형성이 중요한데(Tett, 2004) 학교 전체 접근은
One of the most popular approaches to parameter adaptation in hidden Markov model (HMM) based systems is the maximum likelihood linear regression (MLLR) technique. In this letter, we extend MLLR to factored MLLR (FMLLR) in which the MLLR parameters depend on a continuous-valued control vector. Since it is practically impossible to estimate the MLLR parameters for each control vector separately, we propose a compact parametric form of the MLLR parameters. In the proposed approach, each MLLR param
본 연구는 문화역사활동이론(Cultural Historical Activity Theory: CHAT)의 활동 체계를 분석틀로 삼아 수업 전문성 신장 활동인 수업 장학, 수업 컨설팅, 수업 비평을 살펴보았다. 그 과정에서 각 활동을 구분하는 기준 질문을 도출하였으며 그 질문들을 바탕으로 세 활동의 관계를 통시적인 관점과 공시적인 관점에서 설명하고자 했다. CHAT은 수업 전문성 신장의 단위를 교사 개인이 아니라 수업을 중심으로 대화를 나누는 집단으로 확장하여 보기를 제안한다. 이 관점에서 본다면 각 활동이 이루어지는 공동체 내에서 역할 바꿈 가능성, 인공물과 규칙의 합의 가능성은 수업 전문성 신장 활동의 성격과 의미를 파악하는 데 중요한 기준 질문이 된다. 이러한 기준 질문을 토대로 세 활동의 관계를 통시적으로 보았을 때 각 활동의 등장은 기존의 활동 체계 속의 문제점들을 해결하려는 시도로 볼 수 있으며 우리나라 수업 전문성 신장 활동 체계를 구성하는 각 요소들은 수정되거나 확장되었
We present a novel method to incorporate temporal correlations into a speech recognition system based on conventional hidden Markov models (HMMs). The temporal correlations are considered to be useful for recognition because of the fact that the speech features of the present frame are highly informative about the feature characteristics of neighboring frames. In this paper, by treating these correlations in the form of conditional probability distributions (PDs), we propose a new technique for
We present various methods for estimating a robust output probability distribution (PD) in speech recognition based on the discrete hidden Markov model (HMM). In speech recognition, we encounter the problem of an insufficient amount of training data, which may cause inaccurate modeling of the HMM parameters, especially the output PD's. In this paper, to enhance the robustness of the output PD's with respect to unseen data, we study two approaches: smoothing and tying of the PD's. We introduce a
Recently, the increasing demand for voice-based authentication systems has encouraged researchers to investigate methods for verifying users with short randomized pass-phrases with constrained vocabulary. The conventional i-vector framework, which has been proven to be a state-of-the-art utterance-level feature extraction technique for speaker verification, is not considered to be an optimal method for this task since it is known to suffer from severe performance degradation when dealing with sh
In this letter, we propose a preprocessor that modifies the input signal such that it can be coded more effectively in a low-bit-rate speech coder. Since most of the low-bit-rate speech coders are designed based on the human speech production mechanism, the perceived quality of the speech reconstructed in the decoder degrades seriously if the original input signal deviates from the pure speech. In order to alleviate this problem, we introduce a criterion that compromises the quantization error w
대표 연구 분야
김남수 교수의 연구를 Nubint에서 더 깊이 살펴보세요
이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.