Skip to main content

김민수 교수

Minsu Kim

KAIST · 컴퓨터과학

연구실 소개

김민수 교수의 연구실은 시각적 입모양에서 음성으로의 변환, 즉 립 리딩과 립 투 스피치 합성 기술에 중점을 두고 있습니다. 특히 저자원 언어나 실외 환경 등 제한된 조건에서도 정확한 음성 복원이 가능한 멀티모달 기반의 지능형 음성 인식 및 생성 기술을 개발하고 있습니다. 메모리 기반의 연관 구조와 생성적 적대 신경망을 활용해 시각 정보의 부족함을 보완하고, 음성과 시각의 상호보완적 관계를 효과적으로 모델링하는 데 기여하고 있습니다.

립 리딩립 투 스피치멀티모달 학습메모리 네트워크저자원 언어

연구 현황

논문 수
76
총 인용 수
739
최근 5년 논문
54
주요 분야
컴퓨터과학

연구 성과 추이

표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.

5개년 연도별 논문 게재 수
54총합
2022
2023
2024
2025
2026
5개년 연도별 피인용 수
441총합
20222023202420252026

주요 논문

15
1
논문|인용수 66·2022
Distinguishing Homophenes Using Multi-Head Visual-Audio Memory for Lip Reading
Minsu Kim, Jeong Hun Yeo, Yong Man Ro
Proceedings of the AAAI Conference on Artificial IntelligenceOA

Recognizing speech from silent lip movement, which is called lip reading, is a challenging task due to 1) the inherent information insufficiency of lip movement to fully represent the speech, and 2) the existence of homophenes that have similar lip movement with different pronunciations. In this paper, we try to alleviate the aforementioned two challenges in lip reading by proposing a Multi-head Visual-audio Memory (MVM). Firstly, MVM is trained with audio-visual datasets and remembers audio rep

Signal ProcessingComputer Science
2
논문|인용수 47·2021
Multi-modality Associative Bridging through Memory: Speech Sound Recollected from Face Video
Minsu Kim, Joanna Hong, Se Jin Park, Yong Man Ro
2021 IEEE/CVF International Conference on Computer Vision (ICCV)

In this paper, we introduce a novel audio-visual multi-modal bridging framework that can utilize both audio and visual information, even with uni-modal inputs. We exploit a memory network that stores source (i.e., visual) and target (i.e., audio) modal representations, where source modal representation is what we are given, and target modal representations are what we want to obtain from the memory network. We then construct an associative bridge between source and target memories that considers

Signal ProcessingComputer Science
3
논문|인용수 33·2021
CroMM-VSR: Cross-Modal Memory Augmented Visual Speech Recognition
Minsu Kim, Joanna Hong, Se Jin Park, Yong Man Ro
SJR Q1IEEE Transactions on Multimedia

Visual Speech Recognition (VSR) is a task that recognizes speech from external appearances of the face ( <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">${\it i}.{\it e}.$</tex-math></inline-formula> , lips) into text. Since the information from the visual lip movements is not sufficient to fully represent the speech, VSR is considered as one of the challenging problems. One possible way to resolve this problem

Signal ProcessingComputer Science
4
preprint|인용수 24·2022
Lip to Speech Synthesis with Visual Context Attentional GAN
Minsu Kim, Joanna Hong, Yong Man Ro
arXiv (Cornell University)OA

In this paper, we propose a novel lip-to-speech generative adversarial network, Visual Context Attentional GAN (VCA-GAN), which can jointly model local and global lip movements during speech synthesis. Specifically, the proposed VCA-GAN synthesizes the speech from local lip visual features by finding a mapping function of viseme-to-phoneme, while global visual context is embedded into the intermediate layers of the generator to clarify the ambiguity in the mapping induced by homophene. To achiev

Signal ProcessingComputer Science
5
book chapter|인용수 24·2022
Speaker-Adaptive Lip Reading with User-Dependent Padding
Minsu Kim, Hyunjun Kim, Yong Man Ro
SJR Q2Lecture notes in computer science
Signal ProcessingComputer Science
6
논문|인용수 21·2023
Lip Reading for Low-resource Languages by Learning and Combining General Speech Knowledge and Language-specific Knowledge
Minsu Kim, Jeong Hun Yeo, Jeongsoo Choi, Yong Man Ro

This paper proposes a novel lip reading framework, especially for low-resource languages, which has not been well addressed in the previous literature. Since low-resource languages do not have enough video-text paired data to train the model to have sufficient power to model lip movements and language, it is regarded as challenging to develop lip reading models for low-resource languages. In order to mitigate the challenge, we try to learn general speech knowledge, the ability to model lip movem

Signal ProcessingComputer Science
7
논문|인용수 20·2023
Lip-to-Speech Synthesis in the Wild with Multi-Task Learning
Minsu Kim, Joanna Hong, Yong Man Ro

Recent studies have shown impressive performance in Lip-to-speech synthesis that aims to reconstruct speech from visual information alone. However, they have been suffering from synthesizing accurate speech in the wild, due to insufficient supervision for guiding the model to infer the correct content. Distinct from the previous methods, in this paper, we develop a powerful Lip2Speech method that can reconstruct speech with correct contents from the input lip movements, even in a wild environmen

Signal ProcessingComputer Science
8
논문|인용수 11·2024
Textless Unit-to-Unit Training for Many-to-Many Multilingual Speech-to-Speech Translation
Minsu Kim, Jeongsoo Choi, Dahun Kim, Yong Man Ro
SJR Q1IEEE/ACM Transactions on Audio Speech and Language Processing

This paper proposes a textless training method for many-to-many multilingual speech-to-speech translation that can also benefit the transfer of pre-trained knowledge to text-based systems, text-to-speech synthesis and text-to-speech translation. To this end, we represent multilingual speech with speech units that are the discretized representations of speech features derived from a self-supervised speech model. By treating the speech units as pseudo-text, we can focus on the linguistic content o

Artificial IntelligenceComputer Science
9
preprint|인용수 8·2022
Multi-modality Associative Bridging through Memory: Speech Sound Recollected from Face Video
Minsu Kim, Joanna Hong, Se Jin Park, Yong Man Ro
arXiv (Cornell University)OA

In this paper, we introduce a novel audio-visual multi-modal bridging framework that can utilize both audio and visual information, even with uni-modal inputs. We exploit a memory network that stores source (i.e., visual) and target (i.e., audio) modal representations, where source modal representation is what we are given, and target modal representations are what we want to obtain from the memory network. We then construct an associative bridge between source and target memories that considers

Signal ProcessingComputer Science
10
논문|인용수 6·2024
Towards Practical and Efficient Image-to-Speech Captioning with Vision-Language Pre-Training and Multi-Modal Tokens
Minsu Kim, Jeongsoo Choi, Soumi Maiti, Jeong Hun Yeo, Shinji Watanabe, Yong Man Ro

In this paper, we propose methods to build a powerful and efficient Image-to-Speech captioning (Im2Sp) model. To this end, we start with importing the rich knowledge related to image comprehension and language modeling from a large-scale pre-trained vision-language model into Im2Sp. We set the output of the proposed Im2Sp as discretized speech units, i.e., the quantized speech features of a self-supervised speech model. The speech units mainly contain linguistic information while suppressing oth

Computer Vision and Pattern RecognitionComputer Science
11
preprint|인용수 4·2023
PartMix: Regularization Strategy to Learn Part Discovery for Visible-Infrared Person Re-identification
Minsu Kim, Seungryong Kim, Jungin Park, Seongheon Park, Kwanghoon Sohn
arXiv (Cornell University)OA

Modern data augmentation using a mixture-based technique can regularize the models from overfitting to the training data in various computer vision applications, but a proper data augmentation technique tailored for the part-based Visible-Infrared person Re-IDentification (VI-ReID) models remains unexplored. In this paper, we present a novel data augmentation technique, dubbed PartMix, that synthesizes the augmented samples by mixing the part descriptors across the modalities to improve the perf

Computer Vision and Pattern RecognitionComputer Science
12
preprint|인용수 3·2023
Prompt Tuning of Deep Neural Networks for Speaker-adaptive Visual Speech Recognition
Minsu Kim, Hyung-Il Kim, Yong Man Ro
arXiv (Cornell University)OA

Visual Speech Recognition (VSR) aims to infer speech into text depending on lip movements alone. As it focuses on visual information to model the speech, its performance is inherently sensitive to personal lip appearances and movements, and this makes the VSR models show degraded performance when they are applied to unseen speakers. In this paper, to remedy the performance degradation of the VSR model on unseen speakers, we propose prompt tuning methods of Deep Neural Networks (DNNs) for speaker

Signal ProcessingComputer Science
13
preprint|인용수 3·2022
Speaker-adaptive Lip Reading with User-dependent Padding
Minsu Kim, Hyunjun Kim, Yong Man Ro
arXiv (Cornell University)OA

Lip reading aims to predict speech based on lip movements alone. As it focuses on visual information to model the speech, its performance is inherently sensitive to personal lip appearances and movements. This makes the lip reading models show degraded performance when they are applied to unseen speakers due to the mismatch between training and testing conditions. Speaker adaptation technique aims to reduce this mismatch between train and test speakers, thus guiding a trained model to focus on m

Signal ProcessingComputer Science
14
논문|인용수 3·2024
Efficient Training for Multilingual Visual Speech Recognition: Pre-training with Discretized Visual Speech Representation
Minsu Kim, Jeong Hun Yeo, Se Jin Park, Hyeongseop Rha, Yong Man Ro
OA

This paper explores sentence-level multilingual Visual Speech Recognition (VSR) that can recognize different languages with a single trained model. As the massive multilingual modeling of visual data requires huge computational costs, we propose a novel efficient training strategy, processing with visual speech units. Through analysis, we confirm that the visual speech units mainly contain viseme information while suppressing non-linguistic information. By using the visual speech units as the in

Artificial IntelligenceComputer Science
15
논문|인용수 2·2020
Robust Video Facial Authentication With Unsupervised Mode Disentanglement
Minsu Kim, Hong Joo Lee, Sangmin Lee, Yong Man Ro

Deep learning-based video facial authentication has limitations when it comes to real-world applications, due to large mode variations such as illumination, pose, and eyeglasses variations in real-life situations. Many of existing mode-invariant facial authentication methods need labels of each mode. However, the label information could not be always available in practice. To alleviate this problem, we develop an unsupervised mode disentangling method for video facial authentication. By matching

Computer Vision and Pattern RecognitionComputer Science

대표 연구 분야

Signal ProcessingComputer Vision and Pattern RecognitionArtificial IntelligenceHardware and ArchitectureStatistics and ProbabilityBiomedical Engineering

김민수 교수의 연구를 Nubint에서 더 깊이 살펴보세요

이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.