[Paper Review] VoxSRC 2020: The Second VoxCeleb Speaker Recognition Challenge
The paper describes the second VoxCeleb Speaker Recognition Challenge (VoxSRC2020), covering its tasks (verification and diarisation), new datasets (VoxConverse, VoxMovies), evaluation metrics, baselines, submitted systems, results, and workshop outcomes.
We held the second installment of the VoxCeleb Speaker Recognition Challenge in conjunction with Interspeech 2020. The goal of this challenge was to assess how well current speaker recognition technology is able to diarise and recognize speakers in unconstrained or `in the wild' data. It consisted of: (i) a publicly available speaker recognition and diarisation dataset from YouTube videos together with ground truth annotation and standardised evaluation software; and (ii) a virtual public challenge and workshop held at Interspeech 2020. This paper outlines the challenge, and describes the baselines, methods used, and results. We conclude with a discussion of the progress over the first installment of the challenge.
Motivation & Objective
- Promote and evaluate speaker recognition in unconstrained, real-world conditions ('in the wild').
- Provide public data, evaluation tools, and a public challenge to drive progress in both speaker verification and diarisation.
- Introduce new tasks and metrics to broaden assessment beyond EER, including diarisation measures.
- Offer baselines and analysis to benchmark progress since VoxSRC2019.
Proposed method
- Two tasks: speaker verification (with four tracks) and speaker diarisation (Track 4).
- Public datasets: VoxCeleb variants, VoxMovies for out-of-domain verification, VoxConverse for diarisation.
- New self-supervised track using visual (face) data in training (Track 3).
- Metrics: minDCF and EER for verification; DER and JER for diarisation.
- Baselines: supervised Fast ResNet-34 with mel spectrograms, self-supervised contrastive baseline, and DIHARD-based diarisation baseline.
- Evaluation via CodaLab with time-limited submissions and a workshop (Interspeech 2020).
Experimental results
Research questions
- RQ1How well do state-of-the-art speaker verification and diarisation systems perform under unconstrained, noisy, and cross-domain conditions?
- RQ2Do self-supervised approaches (with or without visual data) approach supervised performance in speaker verification?
- RQ3What is the impact of out-of-domain data (movie material) on verification and diarisation performance?
- RQ4How do diarisation systems handle multispeaker, overlapped conversations in realistic video data?
Key findings
- The top methods for speaker verification across tracks used ECAPA-TDNN and ResNet34 variants with data augmentation and large-margin loss (AAM-softmax).
- In the self-supervised track, performance remained below fully supervised tracks, with EER around 7.21% and minDCF around 0.877 (on the test set).
- VoxMovies-out-of-domain data substantially increased task difficulty, indicating a more challenging test set than VoxCeleb-only data.
- For diarisation (Track 4), the winner achieved DER of 6.23% and JER of 21.52% using conformer-based CSS, Res2Net embeddings, AM-Softmax, and DOVER fusion; second place achieved DER 8.12% and JER 18.35% with VB-HMM post-processing.
- Across verification tracks, the winning submissions significantly outperformed 2019 winners, highlighting substantial progress in one year (e.g., Track 1: 0.177 minDCF, 3.73% EER for the winner).
- The VoxSRC2020 test set was more challenging than VoxSRC2019, as demonstrated by the performance gap when rerunning 2019 winners on the 2020 test set.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.