Skip to main content
QUICK REVIEW

[Paper Review] Visually Informed Binaural Audio Generation without Binaural Audios

Xudong Xu, Hang Zhou|arXiv (Cornell University)|Apr 13, 2021
Speech and Audio Processing47 references4 citations
TL;DR

This paper proposes PseudoBinaural, a method to generate visually informed binaural audio without requiring any recorded binaural data. By leveraging spherical harmonic decomposition and head-related impulse response (HRIR) to map mono audio to binaural outputs, and manually placing visual cues at corresponding source directions, the authors create pseudo visual-stereo pairs for training. The approach achieves competitive performance to supervised methods and boosts state-of-the-art results when combined with real binaural data.

ABSTRACT

Stereophonic audio, especially binaural audio, plays an essential role in immersive viewing environments. Recent research has explored generating visually guided stereophonic audios supervised by multi-channel audio collections. However, due to the requirement of professional recording devices, existing datasets are limited in scale and variety, which impedes the generalization of supervised methods in real-world scenarios. In this work, we propose PseudoBinaural, an effective pipeline that is free of binaural recordings. The key insight is to carefully build pseudo visual-stereo pairs with mono data for training. Specifically, we leverage spherical harmonic decomposition and head-related impulse response (HRIR) to identify the relationship between spatial locations and received binaural audios. Then in the visual modality, corresponding visual cues of the mono data are manually placed at sound source positions to form the pairs. Compared to fully-supervised paradigms, our binaural-recording-free pipeline shows great stability in cross-dataset evaluation and achieves comparable performance under subjective preference. Moreover, combined with binaural recordings, our method is able to further boost the performance of binaural audio generation under supervised settings.

Motivation & Objective

  • To address the scarcity and high cost of binaural audio recordings in training visually informed audio generation models.
  • To overcome the generalization limitations of supervised models trained on controlled, room-specific datasets.
  • To develop a fully unsupervised pipeline that leverages only mono audio and visual data for binaural audio generation.
  • To enable stable, generalizable performance across diverse real-world and in-the-wild scenarios without binaural data.

Proposed method

  • The Mono-Binaural-Mapping procedure uses spherical harmonic decomposition of mono audio to extract zero- and first-order components for spatial audio rendering.
  • Head-related impulse response (HRIR) is applied to convert the decomposed audio components into binaural outputs at any source direction.
  • Visual cues are placed at specific spherical coordinates corresponding to sound source directions using a pre-defined Visual-Coordinate-Mapping.
  • Pseudo visual-stereo pairs are constructed by combining mono audio with visually positioned cues at estimated source locations.
  • The method enables training of mono-to-binaural networks on synthetic pseudo data without requiring real binaural recordings.
  • The framework supports data mixing with varying numbers of sources and incorporates audio-visual source separation as a training auxiliary task.

Experimental results

Research questions

  • RQ1Can binaural audio be effectively generated from mono audio and visual cues without any recorded binaural data?
  • RQ2What is the theoretical relationship between source direction and binaural audio output for a given mono audio?
  • RQ3How can visual cues be systematically mapped to spatial locations to form meaningful pseudo visual-stereo pairs?
  • RQ4Can pseudo data improve generalization and performance in binaural audio generation under both unsupervised and supervised settings?
  • RQ5How does the proposed binaural decoding strategy compare to direct HRIR or ambisonic decoding in terms of perceptual quality?

Key findings

  • PseudoBinaural achieves comparable subjective performance to supervised methods like Mono2Binaural, with users preferring its stereo sensation in 54.7% of evaluations.
  • The method achieves a STFT loss of 0.878 and SNR of 5.316 when using a border azimuth angle of π/3, outperforming other visual field-of-view configurations.
  • User studies show that PseudoBinaural achieves 81% sound localization accuracy, significantly outperforming baseline methods in identifying source positions.
  • The ablation study confirms that combining HRIR and ambisonic decoding yields the best perceptual results, with users rating it highest in 3D hearing experience.
  • When combined with real binaural recordings, PseudoBinaural further boosts performance on standard benchmarks, surpassing current state-of-the-art methods.
  • Visualization analysis confirms that PseudoBinaural focuses attention on actual sound sources, while Mono2Binaural often attends to irrelevant regions like ceilings.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.