Skip to main content

정준성 교수

Joon Son Chung

KAIST 김재철AI대학원 · 컴퓨터과학

연구실 소개

정준성 교수의 연구실은 음성과 시각 정보를 융합한 대규모 자동 코딩 데이터셋을 기반으로, 소음이 많거나 제약이 없는 환경에서도 정확한 화자 인식과 립 리딩 기술을 개발하고 있습니다. 특히, 유튜브 등 오픈소스 미디어에서 자동으로 데이터를 수집·정제하는 파이프라인을 구축하여, 실제 환경에서의 음성 인식 성능을 극대화하는 데 초점을 맞추고 있습니다. 또한, 음성과 영상의 병합 인식, 다화자 음성 분離, 오픈세트 화자 인식 등 실생활 적용에 적합한 지능형 음성 처리 기술을 연구하고 있습니다.

대규모 음성-영상 데이터화자 인식립 리딩소음 환경 인식자연어 음성 처리

연구 현황

논문 수
202
총 인용 수
10,468
최근 5년 논문
137
주요 분야
컴퓨터과학

연구 성과 추이

표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.

5개년 연도별 논문 게재 수
137총합
2022
2023
2024
2025
2026
5개년 연도별 피인용 수
805총합
20222023202420252026

주요 논문

15
1
논문|인용수 2,259·2018
VoxCeleb2: Deep Speaker Recognition
Joon Son Chung, Arsha Nagrani, Andrew Zisserman
OA

<p>The objective of this paper is speaker recognition under noisy and unconstrained conditions.</p> <br/> <p>We make two key contributions. First, we introduce a very large-scale audio-visual speaker recognition dataset collected from open-source media. Using a fully automated pipeline, we curate VoxCeleb2 which contains over a million utterances from over 6,000 speakers. This is several times larger than any publicly available speaker recognition dataset.</p> <b

Artificial IntelligenceComputer Science
2
preprint|인용수 2,077·2017
VoxCeleb: A Large-Scale Speaker Identification Dataset
Arsha Nagrani, Joon Son Chung, Andrew Zisserman
OA

Most existing datasets for speaker identification contain samples obtained under quite constrained conditions, and are usually hand-annotated, hence limited in size. The goal of this paper is to generate a large scale text-independent speaker identification dataset collected 'in the wild'. We make two contributions. First, we propose a fully automated pipeline based on computer vision techniques to create the dataset from open-source media. Our pipeline involves obtaining videos from YouTube; pe

Signal ProcessingComputer Science
3
논문|인용수 761·2018
Deep Audio-Visual Speech Recognition
Triantafyllos Afouras, Joon Son Chung, Andrew Senior, Oriol Vinyals, Andrew Zisserman
SJR Q1IEEE Transactions on Pattern Analysis and Machine IntelligenceOA

The goal of this work is to recognise phrases and sentences being spoken by a talking face, with or without the audio. Unlike previous works that have focussed on recognising a limited number of words or phrases, we tackle lip reading as an open-world problem - unconstrained natural language sentences, and in the wild videos. Our key contributions are: (1) we compare two models for lip reading, one using a CTC loss, and the other using a sequence-to-sequence loss. Both models are built on top of

Signal ProcessingComputer Science
4
논문|인용수 654·2019
Voxceleb: Large-scale speaker verification in the wild
Arsha Nagrani, Joon Son Chung, Weidi Xie, Andrew Zisserman
SJR Q2Computer Speech & LanguageOA

The objective of this work is speaker recognition under noisy and unconstrained conditions. We make two key contributions. First, we introduce a very large-scale audio-visual dataset collected from open source media using a fully automated pipeline. Most existing datasets for speaker identification contain samples obtained under quite constrained conditions, and usually require manual annotations, hence are limited in size. We propose a pipeline based on computer vision techniques to create the

Signal ProcessingComputer Science
5
book chapter|인용수 580·2017
Out of Time: Automated Lip Sync in the Wild
Joon Son Chung, Andrew Zisserman
SJR Q2Lecture notes in computer science
Signal ProcessingComputer Science
6
book chapter|인용수 535·2017
Lip Reading in the Wild
Joon Son Chung, Andrew Zisserman
SJR Q2Lecture notes in computer science
Signal ProcessingComputer Science
7
논문|인용수 394·2020
In Defence of Metric Learning for Speaker Recognition
Joon Son Chung, Jaesung Huh, Seongkyu Mun, Minjae Lee, Hee-Soo Heo, Soyeon Choe, Chiheon Ham, Sunghwan Jung, Bong‐Jin Lee, Icksang Han
OA

The objective of this paper is 'open-set' speaker recognition of unseen speakers, where ideal embeddings should be able to condense information into a compact utterance-level representation that has small intra-speaker and large inter-speaker distance. A popular belief in speaker recognition is that networks trained with classification objectives outperform metric learning methods. In this paper, we present an extensive evaluation of most popular loss functions for speaker recognition on the Vox

Artificial IntelligenceComputer Science
8
논문|인용수 347·2018
The conversation: deep audio-visual speech enhancement
Triantafyllos Afouras, Joon Son Chung, Andrew Zisserman
Oxford University Research Archive (ORA) (University of Oxford)

Our goal is to isolate individual speakers from multi-talker simultaneous speech in videos. Existing works in this area have focussed on trying to separate utterances from known speakers in controlled environments. In this paper, we propose a deep audio-visual speech enhancement network that is able to separate a speaker's voice given lip regions in the corresponding video, by predicting both the magnitude and the phase of the target signal. The method is applicable to speakers unheard and unsee

Signal ProcessingComputer Science
9
논문|인용수 327·2019
Utterance-level Aggregation for Speaker Recognition in the Wild
Weidi Xie, Arsha Nagrani, Joon Son Chung, Andrew Zisserman

The objective of this paper is speaker recognition `in the wild' - where utterances may be of variable length and also contain irrelevant signals. Crucial elements in the design of deep networks for this task are the type of trunk (frame level) network, and the method of temporal aggregation. We propose a powerful speaker recognition deep network, using a `thin-ResNet' trunk architecture, and a dictionary-based NetVLAD or GhostVLAD layer to aggregate features across time, that can be trained end

Artificial IntelligenceComputer Science
10
preprint|인용수 280·2018
LRS3-TED: a large-scale dataset for visual speech recognition
Triantafyllos Afouras, Joon Son Chung, Andrew Zisserman
arXiv (Cornell University)OA

This paper introduces a new multi-modal dataset for visual and audio-visual speech recognition. It includes face tracks from over 400 hours of TED and TEDx videos, along with the corresponding subtitles and word alignment boundaries. The new dataset is substantially larger in scale compared to other public datasets that are available for general research.

Signal ProcessingComputer Science
11
논문|인용수 278·2022
AASIST: Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention Networks
Jee-weon Jung, Hee-Soo Heo, Hemlata Tak, Hye-jin Shim, Joon Son Chung, Bong‐Jin Lee, Ha-Jin Yu, Nicholas Evans
ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

Artefacts that differentiate spoofed from bona-fide utterances can reside in specific temporal or spectral intervals. Their reliable detection usually depends upon computationally demanding ensemble systems where each subsystem is tuned to some specific artefacts. We seek to develop an efficient, single system that can detect a broad range of different spoofing attacks without score-level ensembles. We propose a novel heterogeneous stacking graph attention layer that models artefacts spanning he

Signal ProcessingComputer Science
12
논문|인용수 190·2019
You Said That?: Synthesising Talking Faces from Audio
Amir Jamaludin, Joon Son Chung, Andrew Zisserman
SJR Q1International Journal of Computer VisionOA

We describe a method for generating a video of a talking face. The method takes still images of the target face and an audio speech segment as inputs, and generates a video of the target face lip synched with the audio. The method runs in real time and is applicable to faces and audio not seen at training time. To achieve this we develop an encoder–decoder convolutional neural network (CNN) model that uses a joint embedding of the face and audio to generate synthesised talking face video frames.

Computer Vision and Pattern RecognitionComputer Science
13
논문|인용수 135·2020
Spot the Conversation: Speaker Diarisation in the Wild
Joon Son Chung, Jaesung Huh, Arsha Nagrani, Triantafyllos Afouras, Andrew Zisserman
OA

The goal of this paper is speaker diarisation of videos collected ‘in the wild’.\n<br>\nWe make three key contributions. First, we propose an automatic audio-visual diarisation method for YouTube videos. Our method consists of active speaker detection using audio-visual methods and speaker verification using self-enrolled speaker models. Second, we integrate our method into a semi-automatic dataset creation pipeline which significantly reduces the number of hours required to annotate video

Signal ProcessingComputer Science
14
논문|인용수 94·2018
Learning to lip read words by watching videos
Joon Son Chung, Andrew Zisserman
SJR Q1Computer Vision and Image Understanding
Signal ProcessingComputer Science
15
preprint|인용수 56·2017
You said that
Joon Son Chung, Amir Jamaludin, Andrew Zisserman
arXiv (Cornell University)OA

We present a method for generating a video of a talking face. The method takes as inputs: (i) still images of the target face, and (ii) an audio speech segment; and outputs a video of the target face lip synched with the audio. The method runs in real time and is applicable to faces and audio not seen at training time. To achieve this we propose an encoder-decoder CNN model that uses a joint embedding of the face and audio to generate synthesised talking face video frames. The model is trained

Computer Vision and Pattern RecognitionComputer Science

대표 연구 분야

Artificial IntelligenceSignal ProcessingComputer Vision and Pattern RecognitionHuman-Computer InteractionBiomedical EngineeringManagement Information Systems

정준성 교수의 연구를 Nubint에서 더 깊이 살펴보세요

이 연구실의 논문을 앱에서 열어 AI와 함께 읽고, 핵심을 요약하고, 내 글에 인용하세요.