Skip to main content

김회린 교수

Hoi-Rock Kim

KAIST 전기및전자공학부 · 컴퓨터과학

연구실 소개

김회린 교수의 연구실은 음성인식 및 화자인식 기술의 실용성과 정확도를 높이기 위한 기초 및 응용 연구를 중심으로 전개됩니다. 특히 짧은 음성 문장에서의 화자 식별, 언어 장애를 가진 환자들의 비정상적 발성 패턴을 보완하는 음성 인식 기술, 그리고 잡음 환경에서의 음성 인식 정확도 향상 기법에 중점을 두고 있습니다. 다중 척도 특징 추출, 확률적 히스토그램 보정, 음성 인식 기반의 언어 장애 평가 등 다양한 기술적 접근을 통해 실제 환경에서의 음성 처리 성능을 극대화하고자 합니다.

화자인식비정상 발성음성인식잡음저항성특징추출

연구 현황

논문 수
156
총 인용 수
980
최근 5년 논문
24
주요 분야
컴퓨터과학

연구 성과 추이

표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.

5개년 연도별 논문 게재 수
24총합
2022
2023
2024
2025
2026
5개년 연도별 피인용 수
36총합
20222023202420252026

주요 논문

15
1
논문|인용수 71·2020
Meta-Learning for Short Utterance Speaker Recognition with Imbalance Length Pairs
Seong Min Kye, Youngmoon Jung, Haebeom Lee, Sung Ju Hwang, Hoirin Kim

In practical settings, a speaker recognition system needs to identify a speaker given a short utterance, while the enrollment utterance may be relatively long.However, existing speaker recognition models perform poorly with such short utterances.To solve this problem, we introduce a meta-learning framework for imbalance length pairs.Specifically, we use a Prototypical Networks and train it with a support set of long utterances and a query set of short utterances of varying lengths.Further, since

Artificial IntelligenceComputer Science
2
논문|인용수 56·2017
Regularized Speaker Adaptation of KL-HMM for Dysarthric Speech Recognition
Myungjong Kim, Younggwan Kim, Joohong Yoo, Jun Wang, Hoirin Kim
SJR Q1IEEE Transactions on Neural Systems and Rehabilitation EngineeringOA

This paper addresses the problem of recognizing the speech uttered by patients with dysarthria, which is a motor speech disorder impeding the physical production of speech. Patients with dysarthria have articulatory limitation, and therefore, they often have trouble in pronouncing certain sounds, resulting in undesirable phonetic variation. Modern automatic speech recognition systems designed for regular speakers are ineffective for dysarthric sufferers due to the phonetic variation. To capture

Artificial IntelligenceComputer Science
3
논문|인용수 48·2015
Automatic Intelligibility Assessment of Dysarthric Speech Using Phonologically-Structured Sparse Linear Model
Myung Jong Kim, Younggwan Kim, Hoirin Kim
SJR Q1IEEE/ACM Transactions on Audio Speech and Language Processing

This paper presents a new method for automatically assessing the speech intelligibility of patients with dysarthria, which is a motor speech disorder impeding the physical production of speech. The proposed method consists of two main steps: feature representation and prediction. In the feature representation step, the speech utterance is converted into a phone sequence using an automatic speech recognition technique and is then aligned with a canonical phone sequence from a pronunciation dictio

PhysiologyMedicine
4
논문|인용수 44·2013
Dysarthric speech recognition using dysarthria-severity-dependent and speaker-adaptive models
Myung Jong Kim, Joohong Yoo, Hoirin Kim
Artificial IntelligenceComputer Science
5
논문|인용수 40·2020
Improving Multi-Scale Aggregation Using Feature Pyramid Module for Robust Speaker Verification of Variable-Duration Utterances
Youngmoon Jung, Seong Min Kye, Yeunju Choi, Myunghun Jung, Hoirin Kim
OA

Currently, the most widely used approach for speaker verification is the deep speaker embedding learning. In this approach, we obtain a speaker embedding vector by pooling single-scale features that are extracted from the last layer of a speaker feature extractor. Multi-scale aggregation (MSA), which utilizes multi-scale features from different layers of the feature extractor, has recently been introduced and shows superior performance for variable-duration utterances. To increase the robustness

Artificial IntelligenceComputer Science
6
논문|인용수 38·2006
Frequency-Temporal Filtering for a Robust Audio Fingerprinting Scheme in Real-Noise Environments
Mansoo Park, Hoirin Kim, Seung Hyun Yang
SJR Q2ETRI JournalOA

In a real environment, sound recordings are commonly distorted by channel and background noise, and the performance of audio identification is mainly degraded by them. Recently, Philips introduced a robust and efficient audio fingerprinting scheme applying a differential (high-pass filtering) to the frequency-time sequence of the perceptual filter-bank energies. In practice, however, the robustness of the audio fingerprinting scheme is still important in a real environment. In this letter, we in

Signal ProcessingComputer Science
7
논문|인용수 38·2007
Probabilistic Class Histogram Equalization for Robust Speech Recognition
Young-Joo Suh, Mikyong Ji, Hoirin Kim
SJR Q1IEEE Signal Processing Letters

<para xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> In this letter, a probabilistic class histogram equalization method is proposed to compensate for an acoustic mismatch in noise robust speech recognition. The proposed method aims not only to compensate for the acoustic mismatch between training and test environments but also to reduce the limitations of the conventional histogram equalization. It utilizes multiple class-specific reference and test c

Artificial IntelligenceComputer Science
8
논문|인용수 37·2015
Robust sound event classification using LBP-HOG based bag-of-audio-words feature representation
Hyungjun Lim, Myung Jong Kim, Hoirin Kim
Signal ProcessingComputer Science
9
preprint|인용수 34·2019
Spatial Pyramid Encoding with Convex Length Normalization for Text-Independent Speaker Verification
Youngmoon Jung, Younggwan Kim, Hyungjun Lim, Yeunju Choi, Hoirin Kim
OA

In this paper, we propose a new pooling method called spatial pyramid encoding (SPE) to generate speaker embeddings for text-independent speaker verification. We first partition the output feature maps from a deep residual network (ResNet) into increasingly fine sub-regions and extract speaker embeddings from each sub-region through a learnable dictionary encoding layer. These embeddings are concatenated to obtain the final speaker representation. The SPE layer not only generates a fixed-dimensi

Artificial IntelligenceComputer Science
10
논문|인용수 33·2018
Joint Learning Using Denoising Variational Autoencoders for Voice Activity Detection
Youngmoon Jung, Younggwan Kim, Yeunju Choi, Hoirin Kim
Signal ProcessingComputer Science
11
preprint|인용수 29·2020
Meta-Learned Confidence for Few-shot Learning
Seong Min Kye, Haebeom Lee, Hoirin Kim, Sung Ju Hwang
arXiv (Cornell University)OA

Transductive inference is an effective means of tackling the data deficiency problem in few-shot learning settings. A popular transductive inference technique for few-shot metric-based approaches, is to update the prototype of each class with the mean of the most confident query examples, or confidence-weighted average of all the query samples. However, a caveat here is that the model confidence may be unreliable, which may lead to incorrect predictions. To tackle this issue, we propose to meta-

Artificial IntelligenceComputer Science
12
논문|인용수 28·2012
Multiple Acoustic Model-Based Discriminative Likelihood Ratio Weighting for Voice Activity Detection
Young-Joo Suh, Hoirin Kim
SJR Q1IEEE Signal Processing Letters

In this letter, we propose a novel statistical voice activity detection (VAD) technique. The proposed technique employs probabilistically derived multiple acoustic models to effectively optimize the weights on frequency domain likelihood ratios with the discriminative training approach for more accurate voice activity detection. Experiments performed on various AURORA noisy environments showed that the proposed approach produces meaningful performance improvements over the single acoustic model-

Signal ProcessingComputer Science
13
논문|인용수 28·2020
Deep MOS Predictor for Synthetic Speech Using Cluster-Based Modeling
Yeunju Choi, Youngmoon Jung, Hoirin Kim
OA

While deep learning has made impressive progress in speech synthesis and voice conversion, the assessment of the synthesized speech is still carried out by human participants. Several recent papers have proposed deep-learning-based assessment models and shown the potential to automate the speech quality assessment. To improve the previously proposed assessment model, MOSNet, we propose three models using cluster-based modeling methods: using a global quality token (GQT) layer, using an Encoding

Artificial IntelligenceComputer Science
14
논문|인용수 26·2021
Neural MOS Prediction for Synthesized Speech Using Multi-Task Learning with Spoofing Detection and Spoofing Type Classification
Yeunju Choi, Youngmoon Jung, Hoirin Kim

Several studies have proposed deep-learning-based models to predict the mean opinion score (MOS) of synthesized speech, showing the possibility of replacing human raters. However, inter- and intra-rater variability in MOSs makes it hard to en-sure the high performance of the models. In this paper, we propose a multi-task learning (MTL) method to improve the performance of a MOS prediction model using the following two auxiliary tasks: spoofing detection (SD) and spoofing type classification (STC

Artificial IntelligenceComputer Science
15
논문|인용수 24·2020
Multi-Task Network for Noise-Robust Keyword Spotting and Speaker Verification Using CTC-Based Soft VAD and Global Query Attention
Myunghun Jung, Youngmoon Jung, Jahyun Goo, Hoirin Kim

Keyword spotting (KWS) and speaker verification (SV) have been studied independently although it is known that acoustic and speaker domains are complementary. In this paper, we propose a multi-task network that performs KWS and SV simultaneously to fully utilize the interrelated domain information. The multi-task network tightly combines sub-networks aiming at performance improvement in challenging conditions such as noisy environments, open-vocabulary KWS, and short-duration SV, by introducing

Artificial IntelligenceComputer Science

대표 연구 분야

Artificial IntelligenceSignal ProcessingExperimental and Cognitive PsychologyComputer Networks and CommunicationsComputer Vision and Pattern RecognitionAerospace Engineering

김회린 교수의 연구를 Nubint에서 더 깊이 살펴보세요

이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.