Skip to main content
QUICK REVIEW

[Paper Review] STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events

Kazuki Shimada, Archontis Politis|arXiv (Cornell University)|Jun 15, 2023
Speech and Audio Processing7 citations
TL;DR

This paper introduces STARSS23, a novel audio-visual dataset of real-world soundscapes with multichannel audio, synchronized video, and spatiotemporal annotations of sound events. It proposes an audio-visual sound event localization and detection (SELD) task that leverages visual object positions to improve DOA estimation, demonstrating significant performance gains over audio-only baselines using human-confirmed DOA labels from motion capture data.

ABSTRACT

While direction of arrival (DOA) of sound events is generally estimated from multichannel audio data recorded in a microphone array, sound events usually derive from visually perceptible source objects, e.g., sounds of footsteps come from the feet of a walker. This paper proposes an audio-visual sound event localization and detection (SELD) task, which uses multichannel audio and video information to estimate the temporal activation and DOA of target sound events. Audio-visual SELD systems can detect and localize sound events using signals from a microphone array and audio-visual correspondence. We also introduce an audio-visual dataset, Sony-TAu Realistic Spatial Soundscapes 2023 (STARSS23), which consists of multichannel audio data recorded with a microphone array, video data, and spatiotemporal annotation of sound events. Sound scenes in STARSS23 are recorded with instructions, which guide recording participants to ensure adequate activity and occurrences of sound events. STARSS23 also serves human-annotated temporal activation labels and human-confirmed DOA labels, which are based on tracking results of a motion capture system. Our benchmark results demonstrate the benefits of using visual object positions in audio-visual SELD tasks. The data is available at https://zenodo.org/record/7880637.

Motivation & Objective

  • To address the lack of real-world, multimodal datasets for audio-visual sound event localization and detection (SELD) with accurate spatial and temporal annotations.
  • To investigate how visual modality improves the accuracy of direction-of-arrival (DOA) estimation in complex, natural sound scenes.
  • To provide a benchmark dataset with human-confirmed DOA labels derived from motion capture systems for reliable evaluation of audio-visual SELD systems.
  • To enable research in audio-visual correspondence, source separation, and cross-modal learning using realistic, overlapping, and dynamic sound events.

Proposed method

  • Collecting multichannel audio using a microphone array and synchronized video from 360° cameras in 16 real rooms with 57 participants.
  • Recording sound scenes with guided instructions to ensure diverse, natural, and overlapping sound events such as footsteps, speech, and vacuum cleaners.
  • Using a motion capture system to track physical positions of sound sources and generate ground-truth DOA labels for each sound event.
  • Annotating each frame with temporal activation and DOA for target sound classes, validated by human raters.
  • Creating a benchmark for audio-visual SELD by training models on audio and video inputs to predict spatiotemporal sound event activity and DOA.
  • Releasing the dataset under the MIT license via Zenodo with DOI 10.5281/zenodo.7709051 for public access and extensibility.

Experimental results

Research questions

  • RQ1How does the integration of visual information improve the accuracy of sound event localization and detection in real-world audio-visual scenes?
  • RQ2To what extent do visual object positions reduce ambiguity in DOA estimation for overlapping or transient sound events?
  • RQ3Can audio-visual models trained on STARSS23 outperform audio-only models in complex, dynamic environments with moving sources?
  • RQ4How reliable are human-confirmed DOA labels derived from motion capture systems in real-world soundscapes?
  • RQ5What are the performance gains of audio-visual fusion in SELD tasks compared to audio-only baselines on real, non-synthetic data?

Key findings

  • Audio-visual SELD systems using STARSS23 achieve significantly improved localization accuracy compared to audio-only models, demonstrating the value of visual modality in reducing spatial ambiguity.
  • The dataset contains over 7 hours of real-world recordings with 57 participants across 16 rooms, capturing natural temporal dynamics and spatial movement of sound events.
  • DOA labels are based on motion capture tracking and confirmed by human raters, ensuring high reliability and fidelity to physical source positions.
  • The benchmark results show that visual cues substantially enhance detection and localization performance, especially for transient or overlapping events like footsteps or knocks.
  • STARSS23 supports a wide range of tasks including audio-visual SELD, speaker DOA estimation, sound source localization, and cross-modal retrieval.
  • The dataset is publicly available under the MIT license at Zenodo with DOI 10.5281/zenodo.7709051, enabling reproducible research and community contributions.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.