Skip to main content

Hoi-Rock Kim

Korea Advanced Institute of Science and Technology · 情報科学

研究室紹介

Professor Hoi-Rock Kim's research lab specializes in speech and audio processing, with a strong focus on robust speaker recognition, dysarthric speech analysis, and speech intelligibility assessment. The lab develops advanced machine learning and signal processing techniques to address challenges such as short utterance recognition, phonetic variation in pathological speech, and acoustic mismatches due to noise and channel distortions. Key research directions include meta-learning for imbalanced data, hidden Markov models with Kullback-Leibler divergence for dysarthric speech modeling, and deep learning-based audio fingerprinting for real-world robustness. The lab emphasizes practical applications in healthcare and real-world audio systems.

speaker recognitiondysarthric speechspeech intelligibilityaudio fingerprintingrobust speech processing

Research Overview

Papers
156
Total Citations
980
Papers (5y)
24
Primary Field
情報科学

Research Output Trend

Figures are computed from collected data and may differ slightly.

Publications per year (5y)
24total
2022
2023
2024
2025
2026
Citations per year (5y)
36total
20222023202420252026

Selected Papers

15
1
Article|71 citations·2020
Meta-Learning for Short Utterance Speaker Recognition with Imbalance Length Pairs
Seong Min Kye, Youngmoon Jung, Haebeom Lee, Sung Ju Hwang, Hoirin Kim

In practical settings, a speaker recognition system needs to identify a speaker given a short utterance, while the enrollment utterance may be relatively long.However, existing speaker recognition models perform poorly with such short utterances.To solve this problem, we introduce a meta-learning framework for imbalance length pairs.Specifically, we use a Prototypical Networks and train it with a support set of long utterances and a query set of short utterances of varying lengths.Further, since

Artificial IntelligenceComputer Science
2
Article|56 citations·2017
Regularized Speaker Adaptation of KL-HMM for Dysarthric Speech Recognition
Myungjong Kim, Younggwan Kim, Joohong Yoo, Jun Wang, Hoirin Kim
SJR Q1IEEE Transactions on Neural Systems and Rehabilitation EngineeringOA

This paper addresses the problem of recognizing the speech uttered by patients with dysarthria, which is a motor speech disorder impeding the physical production of speech. Patients with dysarthria have articulatory limitation, and therefore, they often have trouble in pronouncing certain sounds, resulting in undesirable phonetic variation. Modern automatic speech recognition systems designed for regular speakers are ineffective for dysarthric sufferers due to the phonetic variation. To capture

Artificial IntelligenceComputer Science
3
Article|48 citations·2015
Automatic Intelligibility Assessment of Dysarthric Speech Using Phonologically-Structured Sparse Linear Model
Myung Jong Kim, Younggwan Kim, Hoirin Kim
SJR Q1IEEE/ACM Transactions on Audio Speech and Language Processing

This paper presents a new method for automatically assessing the speech intelligibility of patients with dysarthria, which is a motor speech disorder impeding the physical production of speech. The proposed method consists of two main steps: feature representation and prediction. In the feature representation step, the speech utterance is converted into a phone sequence using an automatic speech recognition technique and is then aligned with a canonical phone sequence from a pronunciation dictio

PhysiologyMedicine
4
Article|44 citations·2013
Dysarthric speech recognition using dysarthria-severity-dependent and speaker-adaptive models
Myung Jong Kim, Joohong Yoo, Hoirin Kim
Artificial IntelligenceComputer Science
5
Article|40 citations·2020
Improving Multi-Scale Aggregation Using Feature Pyramid Module for Robust Speaker Verification of Variable-Duration Utterances
Youngmoon Jung, Seong Min Kye, Yeunju Choi, Myunghun Jung, Hoirin Kim
OA

Currently, the most widely used approach for speaker verification is the deep speaker embedding learning. In this approach, we obtain a speaker embedding vector by pooling single-scale features that are extracted from the last layer of a speaker feature extractor. Multi-scale aggregation (MSA), which utilizes multi-scale features from different layers of the feature extractor, has recently been introduced and shows superior performance for variable-duration utterances. To increase the robustness

Artificial IntelligenceComputer Science
6
Article|38 citations·2006
Frequency-Temporal Filtering for a Robust Audio Fingerprinting Scheme in Real-Noise Environments
Mansoo Park, Hoirin Kim, Seung Hyun Yang
SJR Q2ETRI JournalOA

In a real environment, sound recordings are commonly distorted by channel and background noise, and the performance of audio identification is mainly degraded by them. Recently, Philips introduced a robust and efficient audio fingerprinting scheme applying a differential (high-pass filtering) to the frequency-time sequence of the perceptual filter-bank energies. In practice, however, the robustness of the audio fingerprinting scheme is still important in a real environment. In this letter, we in

Signal ProcessingComputer Science
7
Article|38 citations·2007
Probabilistic Class Histogram Equalization for Robust Speech Recognition
Young-Joo Suh, Mikyong Ji, Hoirin Kim
SJR Q1IEEE Signal Processing Letters

<para xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> In this letter, a probabilistic class histogram equalization method is proposed to compensate for an acoustic mismatch in noise robust speech recognition. The proposed method aims not only to compensate for the acoustic mismatch between training and test environments but also to reduce the limitations of the conventional histogram equalization. It utilizes multiple class-specific reference and test c

Artificial IntelligenceComputer Science
8
Article|37 citations·2015
Robust sound event classification using LBP-HOG based bag-of-audio-words feature representation
Hyungjun Lim, Myung Jong Kim, Hoirin Kim
Signal ProcessingComputer Science
9
Preprint|34 citations·2019
Spatial Pyramid Encoding with Convex Length Normalization for Text-Independent Speaker Verification
Youngmoon Jung, Younggwan Kim, Hyungjun Lim, Yeunju Choi, Hoirin Kim
OA

In this paper, we propose a new pooling method called spatial pyramid encoding (SPE) to generate speaker embeddings for text-independent speaker verification. We first partition the output feature maps from a deep residual network (ResNet) into increasingly fine sub-regions and extract speaker embeddings from each sub-region through a learnable dictionary encoding layer. These embeddings are concatenated to obtain the final speaker representation. The SPE layer not only generates a fixed-dimensi

Artificial IntelligenceComputer Science
10
Article|33 citations·2018
Joint Learning Using Denoising Variational Autoencoders for Voice Activity Detection
Youngmoon Jung, Younggwan Kim, Yeunju Choi, Hoirin Kim
Signal ProcessingComputer Science
11
Preprint|29 citations·2020
Meta-Learned Confidence for Few-shot Learning
Seong Min Kye, Haebeom Lee, Hoirin Kim, Sung Ju Hwang
arXiv (Cornell University)OA

Transductive inference is an effective means of tackling the data deficiency problem in few-shot learning settings. A popular transductive inference technique for few-shot metric-based approaches, is to update the prototype of each class with the mean of the most confident query examples, or confidence-weighted average of all the query samples. However, a caveat here is that the model confidence may be unreliable, which may lead to incorrect predictions. To tackle this issue, we propose to meta-

Artificial IntelligenceComputer Science
12
Article|28 citations·2012
Multiple Acoustic Model-Based Discriminative Likelihood Ratio Weighting for Voice Activity Detection
Young-Joo Suh, Hoirin Kim
SJR Q1IEEE Signal Processing Letters

In this letter, we propose a novel statistical voice activity detection (VAD) technique. The proposed technique employs probabilistically derived multiple acoustic models to effectively optimize the weights on frequency domain likelihood ratios with the discriminative training approach for more accurate voice activity detection. Experiments performed on various AURORA noisy environments showed that the proposed approach produces meaningful performance improvements over the single acoustic model-

Signal ProcessingComputer Science
13
Article|28 citations·2020
Deep MOS Predictor for Synthetic Speech Using Cluster-Based Modeling
Yeunju Choi, Youngmoon Jung, Hoirin Kim
OA

While deep learning has made impressive progress in speech synthesis and voice conversion, the assessment of the synthesized speech is still carried out by human participants. Several recent papers have proposed deep-learning-based assessment models and shown the potential to automate the speech quality assessment. To improve the previously proposed assessment model, MOSNet, we propose three models using cluster-based modeling methods: using a global quality token (GQT) layer, using an Encoding

Artificial IntelligenceComputer Science
14
Article|26 citations·2021
Neural MOS Prediction for Synthesized Speech Using Multi-Task Learning with Spoofing Detection and Spoofing Type Classification
Yeunju Choi, Youngmoon Jung, Hoirin Kim

Several studies have proposed deep-learning-based models to predict the mean opinion score (MOS) of synthesized speech, showing the possibility of replacing human raters. However, inter- and intra-rater variability in MOSs makes it hard to en-sure the high performance of the models. In this paper, we propose a multi-task learning (MTL) method to improve the performance of a MOS prediction model using the following two auxiliary tasks: spoofing detection (SD) and spoofing type classification (STC

Artificial IntelligenceComputer Science
15
Article|24 citations·2020
Multi-Task Network for Noise-Robust Keyword Spotting and Speaker Verification Using CTC-Based Soft VAD and Global Query Attention
Myunghun Jung, Youngmoon Jung, Jahyun Goo, Hoirin Kim

Keyword spotting (KWS) and speaker verification (SV) have been studied independently although it is known that acoustic and speaker domains are complementary. In this paper, we propose a multi-task network that performs KWS and SV simultaneously to fully utilize the interrelated domain information. The multi-task network tightly combines sub-networks aiming at performance improvement in challenging conditions such as noisy environments, open-vocabulary KWS, and short-duration SV, by introducing

Artificial IntelligenceComputer Science

Research Areas

Artificial IntelligenceSignal ProcessingExperimental and Cognitive PsychologyComputer Networks and CommunicationsComputer Vision and Pattern RecognitionAerospace Engineering

Hoi-Rock Kimの研究をNubintでさらに深く

この研究室の論文をアプリで開き、AIと共に読み、要約し、引用しましょう。