Skip to main content

Minsu Kim

Korea Advanced Institute of Science and Technology · Computer Science

About the Lab

Professor Minsu Kim's research lab specializes in advanced materials synthesis and environmental monitoring, with a focus on developing high-performance functional materials for sustainable applications and creating high-resolution environmental mapping for public health. The lab investigates metal-organic frameworks for gas separation, mechanisms of soil nitrogen transformations under drying conditions, and the synthesis of nanostructured materials with tailored morphologies. Additionally, the lab explores multi-modal AI frameworks for audio-visual signal processing, particularly in low-information or challenging sensory conditions such as silent lip reading.

nanomaterialsair quality mappingsoil biogeochemistrymulti-modal AImembrane materials

Research Overview

Papers
76
Total Citations
739
Papers (5y)
54
Primary Field
Computer Science

Research Output Trend

Figures are computed from collected data and may differ slightly.

Publications per year (5y)
54total
2022
2023
2024
2025
2026
Citations per year (5y)
441total
20222023202420252026

Selected Papers

15
1
Article|66 citations·2022
Distinguishing Homophenes Using Multi-Head Visual-Audio Memory for Lip Reading
Minsu Kim, Jeong Hun Yeo, Yong Man Ro
Proceedings of the AAAI Conference on Artificial IntelligenceOA

Recognizing speech from silent lip movement, which is called lip reading, is a challenging task due to 1) the inherent information insufficiency of lip movement to fully represent the speech, and 2) the existence of homophenes that have similar lip movement with different pronunciations. In this paper, we try to alleviate the aforementioned two challenges in lip reading by proposing a Multi-head Visual-audio Memory (MVM). Firstly, MVM is trained with audio-visual datasets and remembers audio rep

Signal ProcessingComputer Science
2
Article|47 citations·2021
Multi-modality Associative Bridging through Memory: Speech Sound Recollected from Face Video
Minsu Kim, Joanna Hong, Se Jin Park, Yong Man Ro
2021 IEEE/CVF International Conference on Computer Vision (ICCV)

In this paper, we introduce a novel audio-visual multi-modal bridging framework that can utilize both audio and visual information, even with uni-modal inputs. We exploit a memory network that stores source (i.e., visual) and target (i.e., audio) modal representations, where source modal representation is what we are given, and target modal representations are what we want to obtain from the memory network. We then construct an associative bridge between source and target memories that considers

Signal ProcessingComputer Science
3
Article|33 citations·2021
CroMM-VSR: Cross-Modal Memory Augmented Visual Speech Recognition
Minsu Kim, Joanna Hong, Se Jin Park, Yong Man Ro
SJR Q1IEEE Transactions on Multimedia

Visual Speech Recognition (VSR) is a task that recognizes speech from external appearances of the face ( <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">${\it i}.{\it e}.$</tex-math></inline-formula> , lips) into text. Since the information from the visual lip movements is not sufficient to fully represent the speech, VSR is considered as one of the challenging problems. One possible way to resolve this problem

Signal ProcessingComputer Science
4
Preprint|24 citations·2022
Lip to Speech Synthesis with Visual Context Attentional GAN
Minsu Kim, Joanna Hong, Yong Man Ro
arXiv (Cornell University)OA

In this paper, we propose a novel lip-to-speech generative adversarial network, Visual Context Attentional GAN (VCA-GAN), which can jointly model local and global lip movements during speech synthesis. Specifically, the proposed VCA-GAN synthesizes the speech from local lip visual features by finding a mapping function of viseme-to-phoneme, while global visual context is embedded into the intermediate layers of the generator to clarify the ambiguity in the mapping induced by homophene. To achiev

Signal ProcessingComputer Science
5
Book Chapter|24 citations·2022
Speaker-Adaptive Lip Reading with User-Dependent Padding
Minsu Kim, Hyunjun Kim, Yong Man Ro
SJR Q2Lecture notes in computer science
Signal ProcessingComputer Science
6
Article|21 citations·2023
Lip Reading for Low-resource Languages by Learning and Combining General Speech Knowledge and Language-specific Knowledge
Minsu Kim, Jeong Hun Yeo, Jeongsoo Choi, Yong Man Ro

This paper proposes a novel lip reading framework, especially for low-resource languages, which has not been well addressed in the previous literature. Since low-resource languages do not have enough video-text paired data to train the model to have sufficient power to model lip movements and language, it is regarded as challenging to develop lip reading models for low-resource languages. In order to mitigate the challenge, we try to learn general speech knowledge, the ability to model lip movem

Signal ProcessingComputer Science
7
Article|20 citations·2023
Lip-to-Speech Synthesis in the Wild with Multi-Task Learning
Minsu Kim, Joanna Hong, Yong Man Ro

Recent studies have shown impressive performance in Lip-to-speech synthesis that aims to reconstruct speech from visual information alone. However, they have been suffering from synthesizing accurate speech in the wild, due to insufficient supervision for guiding the model to infer the correct content. Distinct from the previous methods, in this paper, we develop a powerful Lip2Speech method that can reconstruct speech with correct contents from the input lip movements, even in a wild environmen

Signal ProcessingComputer Science
8
Article|11 citations·2024
Textless Unit-to-Unit Training for Many-to-Many Multilingual Speech-to-Speech Translation
Minsu Kim, Jeongsoo Choi, Dahun Kim, Yong Man Ro
SJR Q1IEEE/ACM Transactions on Audio Speech and Language Processing

This paper proposes a textless training method for many-to-many multilingual speech-to-speech translation that can also benefit the transfer of pre-trained knowledge to text-based systems, text-to-speech synthesis and text-to-speech translation. To this end, we represent multilingual speech with speech units that are the discretized representations of speech features derived from a self-supervised speech model. By treating the speech units as pseudo-text, we can focus on the linguistic content o

Artificial IntelligenceComputer Science
9
Preprint|8 citations·2022
Multi-modality Associative Bridging through Memory: Speech Sound Recollected from Face Video
Minsu Kim, Joanna Hong, Se Jin Park, Yong Man Ro
arXiv (Cornell University)OA

In this paper, we introduce a novel audio-visual multi-modal bridging framework that can utilize both audio and visual information, even with uni-modal inputs. We exploit a memory network that stores source (i.e., visual) and target (i.e., audio) modal representations, where source modal representation is what we are given, and target modal representations are what we want to obtain from the memory network. We then construct an associative bridge between source and target memories that considers

Signal ProcessingComputer Science
10
Article|6 citations·2024
Towards Practical and Efficient Image-to-Speech Captioning with Vision-Language Pre-Training and Multi-Modal Tokens
Minsu Kim, Jeongsoo Choi, Soumi Maiti, Jeong Hun Yeo, Shinji Watanabe, Yong Man Ro

In this paper, we propose methods to build a powerful and efficient Image-to-Speech captioning (Im2Sp) model. To this end, we start with importing the rich knowledge related to image comprehension and language modeling from a large-scale pre-trained vision-language model into Im2Sp. We set the output of the proposed Im2Sp as discretized speech units, i.e., the quantized speech features of a self-supervised speech model. The speech units mainly contain linguistic information while suppressing oth

Computer Vision and Pattern RecognitionComputer Science
11
Preprint|4 citations·2023
PartMix: Regularization Strategy to Learn Part Discovery for Visible-Infrared Person Re-identification
Minsu Kim, Seungryong Kim, Jungin Park, Seongheon Park, Kwanghoon Sohn
arXiv (Cornell University)OA

Modern data augmentation using a mixture-based technique can regularize the models from overfitting to the training data in various computer vision applications, but a proper data augmentation technique tailored for the part-based Visible-Infrared person Re-IDentification (VI-ReID) models remains unexplored. In this paper, we present a novel data augmentation technique, dubbed PartMix, that synthesizes the augmented samples by mixing the part descriptors across the modalities to improve the perf

Computer Vision and Pattern RecognitionComputer Science
12
Preprint|3 citations·2023
Prompt Tuning of Deep Neural Networks for Speaker-adaptive Visual Speech Recognition
Minsu Kim, Hyung-Il Kim, Yong Man Ro
arXiv (Cornell University)OA

Visual Speech Recognition (VSR) aims to infer speech into text depending on lip movements alone. As it focuses on visual information to model the speech, its performance is inherently sensitive to personal lip appearances and movements, and this makes the VSR models show degraded performance when they are applied to unseen speakers. In this paper, to remedy the performance degradation of the VSR model on unseen speakers, we propose prompt tuning methods of Deep Neural Networks (DNNs) for speaker

Signal ProcessingComputer Science
13
Preprint|3 citations·2022
Speaker-adaptive Lip Reading with User-dependent Padding
Minsu Kim, Hyunjun Kim, Yong Man Ro
arXiv (Cornell University)OA

Lip reading aims to predict speech based on lip movements alone. As it focuses on visual information to model the speech, its performance is inherently sensitive to personal lip appearances and movements. This makes the lip reading models show degraded performance when they are applied to unseen speakers due to the mismatch between training and testing conditions. Speaker adaptation technique aims to reduce this mismatch between train and test speakers, thus guiding a trained model to focus on m

Signal ProcessingComputer Science
14
Article|3 citations·2024
Efficient Training for Multilingual Visual Speech Recognition: Pre-training with Discretized Visual Speech Representation
Minsu Kim, Jeong Hun Yeo, Se Jin Park, Hyeongseop Rha, Yong Man Ro
OA

This paper explores sentence-level multilingual Visual Speech Recognition (VSR) that can recognize different languages with a single trained model. As the massive multilingual modeling of visual data requires huge computational costs, we propose a novel efficient training strategy, processing with visual speech units. Through analysis, we confirm that the visual speech units mainly contain viseme information while suppressing non-linguistic information. By using the visual speech units as the in

Artificial IntelligenceComputer Science
15
Article|2 citations·2020
Robust Video Facial Authentication With Unsupervised Mode Disentanglement
Minsu Kim, Hong Joo Lee, Sangmin Lee, Yong Man Ro

Deep learning-based video facial authentication has limitations when it comes to real-world applications, due to large mode variations such as illumination, pose, and eyeglasses variations in real-life situations. Many of existing mode-invariant facial authentication methods need labels of each mode. However, the label information could not be always available in practice. To alleviate this problem, we develop an unsupervised mode disentangling method for video facial authentication. By matching

Computer Vision and Pattern RecognitionComputer Science

Research Areas

Signal ProcessingComputer Vision and Pattern RecognitionArtificial IntelligenceHardware and ArchitectureStatistics and ProbabilityBiomedical Engineering

Dive deeper into Minsu Kim's research on Nubint

Open this lab's papers in the app to read with AI, summarize, and cite in your writing.