[Paper Review] SMS-WSJ: Database, performance measures, and baseline recipe for multi-channel source separation and recognition
SMS-WSJ introduces a large multi-channel simulated WSJ-based database with random room geometries and a complete baseline for multi-speaker source separation and recognition, including metrics discussion and Kaldi/TDNN-F baselines.
We present a multi-channel database of overlapping speech for training, evaluation, and detailed analysis of source separation and extraction algorithms: SMS-WSJ -- Spatialized Multi-Speaker Wall Street Journal. It consists of artificially mixed speech taken from the WSJ database, but unlike earlier databases we consider all WSJ0+1 utterances and take care of strictly separating the speaker sets present in the training, validation and test sets. When spatializing the data we ensure a high degree of randomness w.r.t. room size, array center and rotation, as well as speaker position. Furthermore, this paper offers a critical assessment of recently proposed measures of source separation performance. Alongside the code to generate the database we provide a source separation baseline and a Kaldi recipe with competitive word error rates to provide common ground for evaluation.
Motivation & Objective
- Present a multi-channel overlapping speech database built on WSJ utterances with random geometry for controllable realism.
- Critically assess performance measures for multi-channel source separation and recognition.
- Provide a baseline BSS pipeline and an ASR recipe to enable fair comparisons and reproducibility.
Proposed method
- Construct a 33,561 training, 491 validation, and 333 test mixtures using WSJ si284, dev93, and eval92 utterances downsampled to 8 kHz.
- Simulate room impulse responses with random room size, array position, and speaker locations, using a circular 10 cm radius array and random delays to separate early and late speech components.
- Evaluate multiple SDR variants (SI-SDR, BSS-Eval SDR) and perceptual metrics (PESQ, STOI) and WER to provide a comprehensive assessment of separation quality and downstream recognition.
- Provide a source separation baseline based on a complex angular central Gaussian mixture model (cACGMM) with masking and MVDR beamforming, plus distortion masking for covariance estimation.
- Offer a Kaldi-based speech recognition baseline using a TDNN-F acoustic model trained on early-arriving speech images, enabling competitive WER baselines.
- Documentation and code are provided to reproduce the database, metrics, and baselines (SMS-WSJ repository).
Experimental results
Research questions
- RQ1How does multi-channel, far-field speech separation perform under diverse and randomized geometric configurations?
- RQ2What are the most reliable performance metrics for assessing multi-channel BSS in reverberant conditions, and how should they be interpreted?
- RQ3Can a practical baseline BSS pipeline and ASR recipe achieve competitive performance on SMS-WSJ data?
- RQ4How do different baselines (e.g., masking, MVDR, and various beamformers) impact downstream WER in a Kaldi ASR setup?
Key findings
- The SMS-WSJ database provides a large, diverse, and fully reproducible multi-channel dataset based on WSJ utterances with random room geometry and sources, enabling robust evaluation of separation algorithms.
- Multiple SDR variants and perceptual metrics reveal that BSS-Eval SDR with the source signal as reference is stable across channel choices and remains informative for evaluating far-field separation, whereas SI-SDR can be sensitive to short FIR-like distortions.
- The baseline cACGMM with masking and MVDR beamforming improves WER over masking alone, demonstrating the benefit of spatial clustering plus beamforming.
- Using early-arriving speech images for ASR alignment yields favorable acoustic model training in the presence of spatially mixed speech, with the Kaldi TDNN-F recipe achieving competitive WER.
- The authors recommend using multiple complementary metrics (including WER) and favor BSS-Eval SDR with the source signal reference over SI-SDR for far-field evaluation.
- Table 2 indicates that MVDR-based baselines provide better WER than masking alone on the SMS-WSJ test set.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.