Skip to main content
QUICK REVIEW

[Paper Review] Earnings-22: A Practical Benchmark for Accents in the Wild

Miguel del Río, Peter Ha|arXiv (Cornell University)|Mar 29, 2022
Natural Language Processing Techniques4 citations
TL;DR

Earnings-22 introduces a 119-hour, free-to-use benchmark of global English-language earnings calls featuring 7 regional accents to evaluate ASR model performance on real-world accented speech. The study reveals statistically significant WER disparities across accents—especially in Asian, Other Romance, and Spanish/Portuguese regions—highlighting persistent bias in commercial ASR systems despite overall WER improvements.

ABSTRACT

Modern automatic speech recognition (ASR) systems have achieved superhuman Word Error Rate (WER) on many common corpora despite lacking adequate performance on speech in the wild. Beyond that, there is a lack of real-world, accented corpora to properly benchmark academic and commercial models. To ensure this type of speech is represented in ASR benchmarking, we present Earnings-22, a 125 file, 119 hour corpus of English-language earnings calls gathered from global companies. We run a comparison across 4 commercial models showing the variation in performance when taking country of origin into consideration. Looking at hypothesis transcriptions, we explore errors common to all ASR systems tested. By examining Individual Word Error Rate (IWER), we find that key speech features impact model performance more for certain accents than others. Earnings-22 provides a free-to-use benchmark of real-world, accented audio to bridge academic and industrial research.

Motivation & Objective

  • To address the lack of real-world, accented speech corpora for benchmarking ASR systems.
  • To evaluate commercial ASR models on real-world, accented English speech from global companies.
  • To identify systematic performance disparities across regional accents using WER and IWER metrics.
  • To analyze how specific transcript features like filled pauses and word fragments affect ASR error rates across accents.
  • To promote equitable ASR development by exposing bias in current models through a publicly available, diverse dataset.

Proposed method

  • Collected 125 earnings calls from 27 countries, representing 7 regional accent groups based on company headquarters.
  • Produced high-quality verbatim transcripts via human transcription with post-verification and NER tagging.
  • Converted transcripts into a standardized $.nlp$ format with metadata and tokenization.
  • Applied Monte Carlo permutation tests to compare IWER across accent regions and transcript features.
  • Used WER and IWER to measure model performance differences across regions and linguistic features like filled pauses and word fragments.
  • Defined regions using linguistic and geographical groupings, with special attention to underrepresented accents like Nigerian and Ghanaian.

Experimental results

Research questions

  • RQ1How do commercial ASR models perform across different regional accents in real-world earnings calls?
  • RQ2Are performance disparities across accents statistically significant, or due to random variation?
  • RQ3Which transcript features—such as filled pauses or word fragments—most significantly impact ASR error rates in different accents?
  • RQ4To what extent do model errors correlate with specific linguistic or phonetic characteristics of regional accents?
  • RQ5Can a publicly available, diverse corpus of accented speech improve bias detection and mitigation in ASR systems?

Key findings

  • The Asian region showed a statistically significant WER difference from the English region at the 0.005 level, indicating substantial performance degradation.
  • The Other Romance and Spanish/Portuguese regions also showed statistically significant WER differences from the English region at the 0.05 level.
  • Word fragments had a strong impact on IWER across all regions, with p-values of 0.000∗∗ in all cases, indicating high statistical significance.
  • Filled pauses impacted IWER in most regions, contradicting prior work that found no such effect, suggesting regional variation in pause production affects model recognition.
  • In the English and Germanic regions, models performed worse on words preceding filled pauses, while in the Asian region, performance was better on words before filled pauses, indicating region-specific model confusion.
  • The model’s ability to recognize word fragments was consistently poor across all regions, with p-values of 0.000∗∗, showing a persistent error pattern.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.