[Paper Review] A scalable noisy speech dataset and online subjective test framework
This paper introduces MS-SNSD, a scalable noisy speech dataset and an online crowdsourced subjective testing framework for evaluating speech enhancement algorithms. By enabling large-scale subjective MOS evaluations, the authors demonstrate that increasing dataset size improves noise suppression performance, and show that subjective MOS remains essential despite objective metrics like PESQ and POLQA.
Background noise is a major source of quality impairments in Voice over Internet Protocol (VoIP) and Public Switched Telephone Network (PSTN) calls. Recent work shows the efficacy of deep learning for noise suppression, but the datasets have been relatively small compared to those used in other domains (e.g., ImageNet) and the associated evaluations have been more focused. In order to better facilitate deep learning research in Speech Enhancement, we present a noisy speech dataset (MS-SNSD) that can scale to arbitrary sizes depending on the number of speakers, noise types, and Speech to Noise Ratio (SNR) levels desired. We show that increasing dataset sizes increases noise suppression performance as expected. In addition, we provide an open-source evaluation methodology to evaluate the results subjectively at scale using crowdsourcing, with a reference algorithm to normalize the results. To demonstrate the dataset and evaluation framework we apply it to several noise suppressors and compare the subjective Mean Opinion Score (MOS) with objective quality measures such as SNR, PESQ, POLQA, and VISQOL and show why MOS is still required. Our subjective MOS evaluation is the first large scale evaluation of Speech Enhancement algorithms that we are aware of.
Motivation & Objective
- Address the lack of large-scale, diverse noisy speech datasets for deep learning in speech enhancement.
- Overcome limitations in existing subjective evaluation methods, which are time-consuming and not scalable.
- Develop a reproducible, open-source framework for large-scale, crowdsourced subjective testing of speech enhancement algorithms.
- Demonstrate the importance of subjective Mean Opinion Score (MOS) by comparing it with objective metrics like SNR, PESQ, POLQA, and VISQOL.
- Enable standardized, normalized evaluation of speech enhancement models across varying noise types, SNR levels, and speaker populations.
Proposed method
- Design a modular, extensible dataset generation pipeline that supports arbitrary combinations of speakers, noise types, and SNR levels.
- Use a synthetic data generation approach to create a large, diverse, and scalable noisy speech dataset (MS-SNSD) suitable for training and evaluation.
- Implement an online, crowdsourced subjective testing framework using Amazon Mechanical Turk to collect Mean Opinion Scores (MOS) at scale.
- Introduce a reference algorithm-based normalization technique to calibrate subjective scores across different raters and sessions.
- Integrate objective quality metrics (SNR, PESQ, POLQA, VISQOL) alongside subjective MOS for comparative analysis.
- Ensure reproducibility and open access by releasing the dataset and evaluation framework under an open-source license.
Experimental results
Research questions
- RQ1How does increasing the size of a noisy speech dataset impact the performance of deep learning-based noise suppression models?
- RQ2Can a scalable, crowdsourced framework reliably produce consistent and valid subjective Mean Opinion Scores (MOS) for speech enhancement?
- RQ3How do objective speech quality metrics (e.g., PESQ, POLQA, VISQOL) correlate with subjective MOS in real-world noisy conditions?
- RQ4To what extent do existing objective metrics fail to capture perceptual quality improvements that are evident in subjective listening tests?
- RQ5Can a standardized, open-source evaluation framework improve reproducibility and fairness in speech enhancement research?
Key findings
- Increasing the size of the MS-SNSD dataset leads to measurable improvements in noise suppression performance, confirming the benefits of data scale in deep learning.
- The proposed online subjective testing framework enables large-scale, reliable MOS collection with consistent results across multiple raters and conditions.
- Subjective MOS remains a necessary evaluation metric, as it captures perceptual quality improvements not fully reflected in objective metrics like PESQ or SNR.
- The correlation between objective metrics and MOS is moderate at best, highlighting the limitations of relying solely on automated measures.
- The framework successfully normalizes subjective scores using a reference algorithm, enabling fair comparison across different models and test conditions.
- This work presents the first large-scale, open, and reproducible subjective evaluation of speech enhancement algorithms using crowdsourcing.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.