Skip to main content
QUICK REVIEW

[Paper Review] Voices Obscured in Complex Environmental Settings (VOICES) corpus

Colleen Richey, Maria Barrios|arXiv (Cornell University)|Apr 13, 2018
Speech and Audio ProcessingComputer Science5 references43 citations
TL;DR

The VOICES corpus provides open, far-field, multi-microphone speech data in realistic noisy rooms, using LibriSpeech foregrounds with various distractor noises, plus baseline ASR and SID evaluations.

ABSTRACT

This paper introduces the Voices Obscured In Complex Environmental Settings (VOICES) corpus, a freely available dataset under Creative Commons BY 4.0. This dataset will promote speech and signal processing research of speech recorded by far-field microphones in noisy room conditions. Publicly available speech corpora are mostly composed of isolated speech at close-range microphony. A typical approach to better represent realistic scenarios, is to convolve clean speech with noise and simulated room response for model training. Despite these efforts, model performance degrades when tested against uncurated speech in natural conditions. For this corpus, audio was recorded in furnished rooms with background noise played in conjunction with foreground speech selected from the LibriSpeech corpus. Multiple sessions were recorded in each room to accommodate for all foreground speech-background noise combinations. Audio was recorded using twelve microphones placed throughout the room, resulting in 120 hours of audio per microphone. This work is a multi-organizational effort led by SRI International and Lab41 with the intent to push forward state-of-the-art distant microphone approaches in signal processing and speech recognition.

Motivation & Objective

  • Provide a freely accessible distant-microphone speech corpus with realistic room acoustics and background noise to advance speech and signal processing research.
  • Offer multi-room recordings with multiple microphone placements and distractor noises to reflect real-world conditions.
  • Present baseline ASR and speaker identification results to benchmark future models on VOICES data.
  • Promote research in event/background detection, source separation, speech enhancement, and localization in reverberant environments.

Proposed method

  • Record foreground LibriSpeech-derived speech in two furnished rooms with distinct acoustic profiles.
  • Play synchronized distractor noises (TV, music, babble) with foreground speech and record with 12 distant microphones.
  • Provide 120 hours of data per microphone (374,688 audio files) at 48 kHz/24-bit (and 16 kHz options) with source audio and transcriptions.
  • Use rotating foreground speaker to simulate head movement and compare reverberant/ noisy conditions.
  • Annotate with orthographic transcripts and speaker labels; compute basic statistics (SNR, RMS, amplitudes) and provide baseline ASR/SID experiments.
Figure 1: Microphone and loudspeaker configuration (not to scale) used for recording sessions in (a) room 1 (146” x 107”) and (b) room 2 (225” x 158”). The foreground loudspeaker (shown here at its $90^{\circ}$ position), orange rectangle, was placed in a corner of the room, and speakers playing noi
Figure 1: Microphone and loudspeaker configuration (not to scale) used for recording sessions in (a) room 1 (146” x 107”) and (b) room 2 (225” x 158”). The foreground loudspeaker (shown here at its $90^{\circ}$ position), orange rectangle, was placed in a corner of the room, and speakers playing noi

Experimental results

Research questions

  • RQ1How does distant-microphone speech recognition perform under realistic room reverberation and distractor noise when trained on near-field or synthetic data?
  • RQ2What is the impact of microphone distance and room acoustics on ASR and SID performance in the VOICES corpus?
  • RQ3Can open VOICES data support robust acoustic model development for speech/speaker recognition, detection, and enhancement in noisy environments?
  • RQ4What baseline performance do standard ASR and SID systems achieve on VOICES data across noise types and microphone placements?

Key findings

  • ASR WER increases significantly with distance and distractor noise, e.g., 9.3% source vs. 33.0% with babble (Room-1, 90°) for the SRI system.
  • Distance degrades SID: EER rises from 5.72% (source) to 15.1–16.6% (Far microphones in Room-1/Room-2) under no distractors.
  • Distractor noise further degrades SID by about 2–3.5 percentage points in EER depending on noise type (No distractor vs TV/Music/Babble).
  • Average SNR across rooms and conditions shows degradation with distance; room-1 average 22.19 dB and room-2 average 19.50 dB.
  • The VOICES dataset includes 120 hours per microphone, 12 microphones, 374,688 audio files, and aligns with LibriSpeech foregrounds for speaker identification and speech recognition tasks.
  • Baseline ASR results indicate substantial degradation in realistic acoustic environments compared with near-field conditions.
Figure 2: The WER performance is affected by distance from the foreground loudspeaker, as well as room acoustic profile.
Figure 2: The WER performance is affected by distance from the foreground loudspeaker, as well as room acoustic profile.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.