Skip to main content
QUICK REVIEW

[Paper Review] Overview of Tasks and Investigation of Subjective Evaluation Methods in Environmental Sound Synthesis and Conversion

Yuki Okamoto, Keisuke Imoto|arXiv (Cornell University)|Aug 27, 2019
Music and Audio Processing17 references6 citations
TL;DR

This paper reviews environmental sound synthesis and conversion tasks, focusing on statistical generative models like WaveNet, and proposes a multi-faceted subjective evaluation framework. It evaluates synthesized sounds using intelligibility, distinguishability, and naturalness metrics, finding that WaveNet-based synthesis still lags in perceptual quality, especially for complex sounds like electric shavers, and recommends combining multiple evaluation methods for robust assessment.

ABSTRACT

Synthesizing and converting environmental sounds have the potential for many applications such as supporting movie and game production, data augmentation for sound event detection and scene classification. Conventional works on synthesizing and converting environmental sounds are based on a physical modeling or concatenative approach. However, there are a limited number of works that have addressed environmental sound synthesis and conversion with statistical generative models; thus, this research area is not yet well organized. In this paper, we review problem definitions, applications, and evaluation methods of environmental sound synthesis and conversion. We then report on environmental sound synthesis using sound event labels, in which we focus on the current performance of statistical environmental sound synthesis and investigate how we should conduct subjective experiments on environmental sound synthesis.

Motivation & Objective

  • To organize and clarify the problem definitions and applications of environmental sound synthesis and conversion, particularly in the context of deep learning.
  • To address the lack of standardized evaluation methods for environmental sound synthesis, especially subjective evaluation approaches.
  • To investigate the current performance of statistical generative models—specifically WaveNet—in synthesizing environmental sounds.
  • To propose a multi-dimensional evaluation framework that includes intelligibility, distinguishability, and naturalness for assessing synthesized environmental sounds.
  • To guide future research by identifying gaps in evaluation practices and recommending best practices for subjective testing in environmental sound synthesis.

Proposed method

  • Conducted subjective evaluation experiments using 24 listeners in a controlled environment with high-fidelity audio equipment (Roland UA-55 interface and SONY MDR-CD900ST headphones).
  • Performed three distinct subjective tests: (1) intelligibility via sound event classification (recall-based), (2) distinguishability via AB preference testing, and (3) naturalness via 5-point Mean Opinion Score (MOS).
  • Used WaveNet-based models to synthesize environmental sounds from sound event or scene labels, focusing on realism and perceptual fidelity.
  • Collected and analyzed data using standard psychophysical testing protocols, including random presentation of real and synthesized sounds.
  • Compared synthesized sound performance against real environmental sounds across multiple metrics to assess perceptual quality and realism.
  • Employed spectrogram analysis to correlate perceptual differences with spectral characteristics, particularly for misclassified or low-naturalness sounds.

Experimental results

Research questions

  • RQ1How well can WaveNet-based models synthesize environmental sounds that are perceptually recognizable as specific sound events?
  • RQ2To what extent can listeners distinguish between real and synthesized environmental sounds in a preference test?
  • RQ3How do listeners rate the naturalness of synthesized environmental sounds compared to real ones using a 5-point MOS scale?
  • RQ4What are the limitations of current statistical generative models (e.g., WaveNet) in reproducing fine spectral structures of complex environmental sounds?
  • RQ5Which combination of subjective evaluation metrics provides the most informative assessment of environmental sound synthesis quality?

Key findings

  • The average F-score for sound event classification (intelligibility) was 86.22% for real sounds and 76.30% for synthesized sounds, indicating that synthesized drum sounds were well recognized, while cup clinking and electric shaver sounds were frequently misclassified.
  • Listeners correctly identified real sounds in a distinguishability AB test with 82.71% accuracy, showing that synthesized sounds from WaveNet are still perceptually distinguishable from real ones.
  • The Mean Opinion Score (MOS) for naturalness showed that synthesized coffee grinder, clock, and maracas sounds scored similarly to real ones, but electric shaver and trash box banging sounds had significantly lower scores.
  • Spectrogram analysis revealed that synthesized electric shaver sounds lacked fine spectral structures, leading to confusion with similar sounds like tearing paper.
  • The study found that intelligibility alone is insufficient for evaluating synthesis quality, as some sounds were correctly classified but rated as unnatural, highlighting the need for multi-metric evaluation.
  • The results support the need to evaluate environmental sound synthesis not only for intelligibility but also for distinguishability and naturalness to ensure perceptual realism.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.