Skip to main content
QUICK REVIEW

[Paper Review] Multilingual Dysarthric Speech Assessment Using Universal Phone Recognition and Language-Specific Phonemic Contrast Modeling

Eunjung Yeo, Julie Liss|arXiv (Cornell University)|Jan 29, 2026
Voice and Speech Disorders0 citations
TL;DR

The paper presents a multilingual phoneme-production assessment framework for dysarthric speech that uses a Universal Phone Recognizer and language-specific phonemic contrasts to compute PER, PFER, and PhonCov, improving correlations with clinician intelligibility ratings across four languages.

ABSTRACT

The growing prevalence of neurological disorders associated with dysarthria motivates the need for automated intelligibility assessment methods that are applicalbe across languages. However, most existing approaches are either limited to a single language or fail to capture language-specific factors shaping intelligibility. We present a multilingual phoneme-production assessment framework that integrates universal phone recognition with language-specific phoneme interpretation using contrastive phonological feature distances for phone-to-phoneme mapping and sequence alignment. The framework yields three metrics: phoneme error rate (PER), phonological feature error rate (PFER), and a newly proposed alignment-free measure, phoneme coverage (PhonCov). Analysis on English, Spanish, Italian, and Tamil show that PER benefits from the combination of mapping and alignment, PFER from alignment alone, and PhonCov from mapping. Further analyses demonstrate that the proposed framework captures clinically meaningful patterns of intelligibility degradation consistent with established observations of dysarthric speech.

Motivation & Objective

  • Motivate scalable, cross-language dysarthria assessment that preserves language-specific intelligibility factors.
  • Integrate universal phone recognition with language-specific phoneme interpretation to produce interpretable metrics.
  • Evaluate how mapping and alignment contribute to metric performance across languages (English, Spanish, Italian, Tamil).
  • Introduce PhonCov as an alignment-free measure of phoneme coverage to complement existing metrics.

Proposed method

  • Use a Universal Phone Recognizer (UPR) to transcribe speech into language-agnostic IPA sequences.
  • Map UPR outputs to each language's phoneme inventory using contrastive phonological feature distances.
  • Apply a weighted Needleman–Wunsch alignment with contrast-aware substitution costs for reference vs. predicted sequences.
  • Compute three metrics: Per for phoneme-level errors, PFER for feature-level differences, and PhonCov for phoneme inventory coverage.
  • Evaluate against clinician intelligibility ratings using Kendall’s tau and bootstrap tests for significance.

Experimental results

Research questions

  • RQ1How does language-specific phoneme interpretation affect the correlation between phoneme-production metrics and intelligibility scores across languages?
  • RQ2What are the relative contributions of phone-to-phoneme mapping and alignment to the performance of PER, PFER, and PhonCov?
  • RQ3Does the alignment-free PhonCov metric provide competitive predictive value compared to alignment-based metrics?
  • RQ4How robust are the results to different universal phone recognizers across languages?
  • RQ5Can a training-free, multilingual phoneme-production framework capture clinically meaningful patterns in dysarthric speech across English, Spanish, Italian, and Tamil?

Key findings

  • Incorporating language-specific processing generally improves correlations with intelligibility across languages.
  • PER benefits most from combining mapping and alignment; PFER benefits mainly from alignment; PhonCov benefits from mapping and remains competitive as an alignment-free measure.
  • PhonCov provides comparable correlations to alignment-based metrics despite not requiring alignment.
  • English shows limited gains from language-specific processing due to near-phoneme-level readiness of UPR outputs.
  • Across languages, no single UPR architecture dominates; language-specific interpretation improves stability of intelligibility prediction.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.