[Paper Review] Quantifying Bias in Automatic Speech Recognition
This paper systematically quantifies bias in a Dutch state-of-the-art ASR system across gender, age, regional accents, and non-native accents, using WER and phoneme-level analyses to identify where bias occurs and propose mitigation strategies.
Automatic speech recognition (ASR) systems promise to deliver objective interpretation of human speech. Practice and recent evidence suggests that the state-of-the-art (SotA) ASRs struggle with the large variation in speech due to e.g., gender, age, speech impairment, race, and accents. Many factors can cause the bias of an ASR system. Our overarching goal is to uncover bias in ASR systems to work towards proactive bias mitigation in ASR. This paper is a first step towards this goal and systematically quantifies the bias of a Dutch SotA ASR system against gender, age, regional accents and non-native accents. Word error rates are compared, and an in-depth phoneme-level error analysis is conducted to understand where bias is occurring. We primarily focus on bias due to articulation differences in the dataset. Based on our findings, we suggest bias mitigation strategies for ASR development.
Motivation & Objective
- Motivate the need to uncover bias in ASR systems and move toward proactive mitigation.
- Quantify bias in a standard Dutch DNN-HMM ASR across gender, age groups, regional accents, and non-native accents.
- Compare word error rates (WERs) and perform phoneme-level error analysis to identify bias sources.
- Provide data-driven bias mitigation suggestions based on empirical findings.
Proposed method
- Use a hybrid DNN-HMM Dutch ASR (TDNN-BLSTM) with LF-MMI training in Kaldi.
- Train on the Dutch CGN corpus and evaluate on Jasmin-CGN extensions to cover gender, age, regional, and non-native accents.
- Compare WERs for read speech and human-machine interaction (HMI) speech separately.
- Convert transcripts to phoneme sequences via the Dutch lexicon and compute phoneme error rate (PER) using Levenshtein alignment.
- Perform phoneme-level analysis to identify which phonemes are most misrecognized across groups.
Experimental results
Research questions
- RQ1How does ASR performance (WER) differ across gender, age groups, regional accents, and non-native accents in Dutch?
- RQ2Does speaking style (read vs. HMI) affect the magnitude of bias in ASR performance?
- RQ3Which phonemes are most frequently misrecognized for different speaker groups, and what does this imply about articulation-related bias?
- RQ4What mitigation strategies can be inferred to reduce bias in Dutch ASR based on the observed results?
Key findings
- Female speech is better recognised than male speech across groups and styles.
- Native Dutch speech is recognised more accurately than non-native speech, with non-native groups showing the largest performance gaps.
- Children and, especially, older adults (65+) exhibit higher WERs, with older adults showing the highest variability and worst performance in some regions.
- Read speech generally yields lower WER than HMI speech, with the gap averaging around 13.7 percentage points for natives and 5.5 points for non-natives.
- Regional accents matter: Flemish Dutch (FL) performs worst among native groups, and region S often shows the strongest bias in HMI speech; older Dutch speakers show stronger regional effects.
- Phoneme-level analysis reveals vowels /œy/, /Y/, /y/, /ø:/ and language-specific realizations as frequent sources of misrecognition across groups; native vs. non-native and regional differences drive distinct error patterns.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.