[Paper Review] The VOiCES from a Distance Challenge 2019 Evaluation Plan
The VOiCES from a Distance Challenge 2019 evaluation plan establishes tasks for speaker recognition and ASR on distant/noisy audio, with fixed/open training conditions, development/evaluation sets, specific primary and llr-based metrics, and submission rules for Interspeech 2019 Special Session.
The "VOiCES from a Distance Challenge 2019" is designed to foster research in the area of speaker recognition and automatic speech recognition (ASR) with the special focus on single channel distant/far-field audio, under noisy conditions. The main objectives of this challenge are to: (i) benchmark state-of-the-art technology in the area of speaker recognition and automatic speech recognition (ASR), (ii) support the development of new ideas and technologies in speaker recognition and ASR, (iii) support new research groups entering the field of distant/far-field speech processing, and (iv) provide a new, publicly available dataset to the community that exhibits realistic distance characteristics.
Motivation & Objective
- Foster progress in distant/far-field speaker recognition and ASR under noisy environments.
- Benchmark state-of-the-art techniques using the VOiCES corpus with realistic reverberation and background noise.
- Provide a publicly available dataset and a framework to compare systems fairly across fixed/open training conditions.
- Encourage new researchers and groups to participate and contribute descriptions and analyses for publication.
- Deliver evaluation data releases (phase 2) and organize a Special Session at Interspeech 2019.
Proposed method
- Define two tasks: speaker recognition and automatic speech recognition (ASR).
- Specify training conditions: fixed (limited public data) and open (any data) for each task.
- Provide development and evaluation data from the VOiCES corpus with reverberation and noise."
- Use primary detection-cost metric C_det and an alternate C_llr for speaker recognition evaluation.
- Score submissions using per-trial LLRs for speaker recognition and WER for ASR, with standardized scoring scripts.
- Require CTM-formatted ASR transcripts and LLR-based score files for speaker recognition submissions.
Experimental results
Research questions
- RQ1How well do state-of-the-art systems perform on distant/far-field speech with real reverberation and background noise?
- RQ2What is the impact of training data restrictions (fixed vs open) on speaker recognition and ASR performance?
- RQ3How do calibration metrics (C_llr) compare across operating points for speaker recognition?
- RQ4What can the VOiCES dataset reveal about system robustness to microphone, room, and distractor variations?
Key findings
- The plan introduces two tasks (speaker recognition and ASR) with fixed and open training conditions to benchmark systems.
- It adopts a NIST SRE-like primary metric for speaker recognition (C_det) and an llr-based alternative (C_llr) for calibration analysis.
- ASR performance is evaluated with Word-Error Rate (WER) using SCTK scoring, mirroring NIST OPENSAT-17 evaluation.
- Phase 2 data expands VOiCES with over 310k audio files across diverse reverberant environments.
- Participants must submit system outputs per condition with standardized naming and CTM/LLR formats, and provide system descriptions for conference publication.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.