Skip to main content
QUICK REVIEW

[Paper Review] EasyCom: An Augmented Reality Dataset to Support Algorithms for Easy Communication in Noisy Environments

Jacob Donley, Vladimir Tourbabin|arXiv (Cornell University)|Jul 9, 2021
Speech and Audio ProcessingComputer Science28 references32 citations
TL;DR

This paper introduces EasyCom, a large egocentric AR dataset with synchronized multi-channel audio, video, and annotations in noisy social settings, and provides a baseline beamforming approach and evaluation.

ABSTRACT

Augmented Reality (AR) as a platform has the potential to facilitate the reduction of the cocktail party effect. Future AR headsets could potentially leverage information from an array of sensors spanning many different modalities. Training and testing signal processing and machine learning algorithms on tasks such as beam-forming and speech enhancement require high quality representative data. To the best of the author's knowledge, as of publication there are no available datasets that contain synchronized egocentric multi-channel audio and video with dynamic movement and conversations in a noisy environment. In this work, we describe, evaluate and release a dataset that contains over 5 hours of multi-modal data useful for training and testing algorithms for the application of improving conversations for an AR glasses wearer. We provide speech intelligibility, quality and signal-to-noise ratio improvement results for a baseline method and show improvements across all tested metrics. The dataset we are releasing contains AR glasses egocentric multi-channel microphone array audio, wide field-of-view RGB video, speech source pose, headset microphone audio, annotated voice activity, speech transcriptions, head bounding boxes, target of speech and source identification labels. We have created and are releasing this dataset to facilitate research in multi-modal AR solutions to the cocktail party problem.

Motivation & Objective

  • Motivate the need for realistic, egocentric multi-modal data to mitigate the cocktail party effect in AR.
  • Describe the EasyCom dataset, including sensors, annotations, and acquisition protocol.
  • Release the dataset publicly to enable research on multi-modal AR solutions for speaking in noisy environments.
  • Provide a baseline signal processing method and quantitative benchmarks to evaluate target speech enhancement in AR setups.

Proposed method

  • Describe the data collection setup in a restaurant-like room with 6m x 7m x 3m dimensions and 12 sessions totaling ~5h of data.
  • Record egocentric multi-channel microphone audio and wide-FOV video from an AR glasses wearer plus headset microphones and motion tracking.
  • Annotate voice activity, speech transcriptions, target of speech, and face/head bounding boxes using human raters and automated tools.
  • Provide calibration data and pose/trajectory information for sensor fusion research.
  • Propose a baseline multi-channel beamformer (maximum DI beamformer) to enhance a target speech source while suppressing noise and distortion.
  • Outline the calculation of beamformer weights using estimated d(omega) and R(omega) and a WOLA processing pipeline.

Experimental results

Research questions

  • RQ1How can multi-modal AR data be used to alleviate the cocktail party problem for AR wearers?
  • RQ2What performance gains can be achieved with a baseline maximum DI beamformer on the EasyCom data across SNR, intelligibility, and quality metrics?
  • RQ3How do egocentric dynamics and multi-sensor fusion affect speech enhancement in realistic noisy environments?

Key findings

  • The dataset contains ~5h18m of data across 323 one-minute segments (approx. 79 GB, CC-BY-NC-4.0).
  • Baseline maximum DI beamformer improves several metrics over the reference microphone under noise and interference conditions (e.g., SNR: -9.27 to -6.62; SegSNR: -14.2 to -10.7; SDR: -8.98 to -7.79; STOI: 0.504 to 0.590; PESQ: 1.17 to 1.27; HASQI: 0.268 to 0.319).
  • Under noise+interferer conditions, the baseline improves SNR from -13.3 to -10.1 and maintains higher STOI/SDR-related metrics compared to the reference mic.
  • The dataset enables evaluation across a broad set of objectives including VO activity detection, speaker diarization, ASR, and audio-visual speech processing, with rich annotations and pose data.
  • The baseline results indicate real-time capable beamforming when leveraging ATFs and relative geometry information from AR wearers.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.