Skip to main content

Joon Son Chung

Korea Advanced Institute of Science and Technology · 情報科学

研究室紹介

Professor Joon Son Chung's research lab specializes in audio-visual speech processing and speaker recognition, focusing on real-world, unconstrained environments. The lab develops large-scale, automatically curated datasets—such as VoxCeleb2—using computer vision and deep learning to enable robust speaker identification and verification. Key research directions include end-to-end lip reading, audio-visual speech separation, and metric learning for open-set speaker recognition, with strong emphasis on generalization to noisy and unseen conditions. The lab also investigates the synergy between visual and audio modalities to improve speech recognition in challenging environments.

speaker recognitionaudio-visual learninglip readinglarge-scale datasetsspeech separation

Research Overview

Papers
202
Total Citations
10,468
Papers (5y)
137
Primary Field
情報科学

Research Output Trend

Figures are computed from collected data and may differ slightly.

Publications per year (5y)
137total
2022
2023
2024
2025
2026
Citations per year (5y)
805total
20222023202420252026

Selected Papers

15
1
Article|2,259 citations·2018
VoxCeleb2: Deep Speaker Recognition
Joon Son Chung, Arsha Nagrani, Andrew Zisserman
OA

<p>The objective of this paper is speaker recognition under noisy and unconstrained conditions.</p> <br/> <p>We make two key contributions. First, we introduce a very large-scale audio-visual speaker recognition dataset collected from open-source media. Using a fully automated pipeline, we curate VoxCeleb2 which contains over a million utterances from over 6,000 speakers. This is several times larger than any publicly available speaker recognition dataset.</p> <b

Artificial IntelligenceComputer Science
2
Preprint|2,077 citations·2017
VoxCeleb: A Large-Scale Speaker Identification Dataset
Arsha Nagrani, Joon Son Chung, Andrew Zisserman
OA

Most existing datasets for speaker identification contain samples obtained under quite constrained conditions, and are usually hand-annotated, hence limited in size. The goal of this paper is to generate a large scale text-independent speaker identification dataset collected 'in the wild'. We make two contributions. First, we propose a fully automated pipeline based on computer vision techniques to create the dataset from open-source media. Our pipeline involves obtaining videos from YouTube; pe

Signal ProcessingComputer Science
3
Article|761 citations·2018
Deep Audio-Visual Speech Recognition
Triantafyllos Afouras, Joon Son Chung, Andrew Senior, Oriol Vinyals, Andrew Zisserman
SJR Q1IEEE Transactions on Pattern Analysis and Machine IntelligenceOA

The goal of this work is to recognise phrases and sentences being spoken by a talking face, with or without the audio. Unlike previous works that have focussed on recognising a limited number of words or phrases, we tackle lip reading as an open-world problem - unconstrained natural language sentences, and in the wild videos. Our key contributions are: (1) we compare two models for lip reading, one using a CTC loss, and the other using a sequence-to-sequence loss. Both models are built on top of

Signal ProcessingComputer Science
4
Article|654 citations·2019
Voxceleb: Large-scale speaker verification in the wild
Arsha Nagrani, Joon Son Chung, Weidi Xie, Andrew Zisserman
SJR Q2Computer Speech & LanguageOA

The objective of this work is speaker recognition under noisy and unconstrained conditions. We make two key contributions. First, we introduce a very large-scale audio-visual dataset collected from open source media using a fully automated pipeline. Most existing datasets for speaker identification contain samples obtained under quite constrained conditions, and usually require manual annotations, hence are limited in size. We propose a pipeline based on computer vision techniques to create the

Signal ProcessingComputer Science
5
Book Chapter|580 citations·2017
Out of Time: Automated Lip Sync in the Wild
Joon Son Chung, Andrew Zisserman
SJR Q2Lecture notes in computer science
Signal ProcessingComputer Science
6
Book Chapter|535 citations·2017
Lip Reading in the Wild
Joon Son Chung, Andrew Zisserman
SJR Q2Lecture notes in computer science
Signal ProcessingComputer Science
7
Article|394 citations·2020
In Defence of Metric Learning for Speaker Recognition
Joon Son Chung, Jaesung Huh, Seongkyu Mun, Minjae Lee, Hee-Soo Heo, Soyeon Choe, Chiheon Ham, Sunghwan Jung, Bong‐Jin Lee, Icksang Han
OA

The objective of this paper is 'open-set' speaker recognition of unseen speakers, where ideal embeddings should be able to condense information into a compact utterance-level representation that has small intra-speaker and large inter-speaker distance. A popular belief in speaker recognition is that networks trained with classification objectives outperform metric learning methods. In this paper, we present an extensive evaluation of most popular loss functions for speaker recognition on the Vox

Artificial IntelligenceComputer Science
8
Article|347 citations·2018
The conversation: deep audio-visual speech enhancement
Triantafyllos Afouras, Joon Son Chung, Andrew Zisserman
Oxford University Research Archive (ORA) (University of Oxford)

Our goal is to isolate individual speakers from multi-talker simultaneous speech in videos. Existing works in this area have focussed on trying to separate utterances from known speakers in controlled environments. In this paper, we propose a deep audio-visual speech enhancement network that is able to separate a speaker's voice given lip regions in the corresponding video, by predicting both the magnitude and the phase of the target signal. The method is applicable to speakers unheard and unsee

Signal ProcessingComputer Science
9
Article|327 citations·2019
Utterance-level Aggregation for Speaker Recognition in the Wild
Weidi Xie, Arsha Nagrani, Joon Son Chung, Andrew Zisserman

The objective of this paper is speaker recognition `in the wild' - where utterances may be of variable length and also contain irrelevant signals. Crucial elements in the design of deep networks for this task are the type of trunk (frame level) network, and the method of temporal aggregation. We propose a powerful speaker recognition deep network, using a `thin-ResNet' trunk architecture, and a dictionary-based NetVLAD or GhostVLAD layer to aggregate features across time, that can be trained end

Artificial IntelligenceComputer Science
10
Preprint|280 citations·2018
LRS3-TED: a large-scale dataset for visual speech recognition
Triantafyllos Afouras, Joon Son Chung, Andrew Zisserman
arXiv (Cornell University)OA

This paper introduces a new multi-modal dataset for visual and audio-visual speech recognition. It includes face tracks from over 400 hours of TED and TEDx videos, along with the corresponding subtitles and word alignment boundaries. The new dataset is substantially larger in scale compared to other public datasets that are available for general research.

Signal ProcessingComputer Science
11
Article|278 citations·2022
AASIST: Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention Networks
Jee-weon Jung, Hee-Soo Heo, Hemlata Tak, Hye-jin Shim, Joon Son Chung, Bong‐Jin Lee, Ha-Jin Yu, Nicholas Evans
ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

Artefacts that differentiate spoofed from bona-fide utterances can reside in specific temporal or spectral intervals. Their reliable detection usually depends upon computationally demanding ensemble systems where each subsystem is tuned to some specific artefacts. We seek to develop an efficient, single system that can detect a broad range of different spoofing attacks without score-level ensembles. We propose a novel heterogeneous stacking graph attention layer that models artefacts spanning he

Signal ProcessingComputer Science
12
Article|190 citations·2019
You Said That?: Synthesising Talking Faces from Audio
Amir Jamaludin, Joon Son Chung, Andrew Zisserman
SJR Q1International Journal of Computer VisionOA

We describe a method for generating a video of a talking face. The method takes still images of the target face and an audio speech segment as inputs, and generates a video of the target face lip synched with the audio. The method runs in real time and is applicable to faces and audio not seen at training time. To achieve this we develop an encoder–decoder convolutional neural network (CNN) model that uses a joint embedding of the face and audio to generate synthesised talking face video frames.

Computer Vision and Pattern RecognitionComputer Science
13
Article|135 citations·2020
Spot the Conversation: Speaker Diarisation in the Wild
Joon Son Chung, Jaesung Huh, Arsha Nagrani, Triantafyllos Afouras, Andrew Zisserman
OA

The goal of this paper is speaker diarisation of videos collected ‘in the wild’.\n<br>\nWe make three key contributions. First, we propose an automatic audio-visual diarisation method for YouTube videos. Our method consists of active speaker detection using audio-visual methods and speaker verification using self-enrolled speaker models. Second, we integrate our method into a semi-automatic dataset creation pipeline which significantly reduces the number of hours required to annotate video

Signal ProcessingComputer Science
14
Article|94 citations·2018
Learning to lip read words by watching videos
Joon Son Chung, Andrew Zisserman
SJR Q1Computer Vision and Image Understanding
Signal ProcessingComputer Science
15
Preprint|56 citations·2017
You said that
Joon Son Chung, Amir Jamaludin, Andrew Zisserman
arXiv (Cornell University)OA

We present a method for generating a video of a talking face. The method takes as inputs: (i) still images of the target face, and (ii) an audio speech segment; and outputs a video of the target face lip synched with the audio. The method runs in real time and is applicable to faces and audio not seen at training time. To achieve this we propose an encoder-decoder CNN model that uses a joint embedding of the face and audio to generate synthesised talking face video frames. The model is trained

Computer Vision and Pattern RecognitionComputer Science

Research Areas

Artificial IntelligenceSignal ProcessingComputer Vision and Pattern RecognitionHuman-Computer InteractionBiomedical EngineeringManagement Information Systems

Joon Son Chungの研究をNubintでさらに深く

この研究室の論文をアプリで開き、AIと共に読み、要約し、引用しましょう。