Skip to main content
QUICK REVIEW

[论文解读] Experiments of ASR-based mispronunciation detection for children and adult English learners

Nina Hosseini-Kivanani, Roberto Gretter|arXiv (Cornell University)|Apr 13, 2021
Speech Recognition and Synthesis参考文献 20被引用 5
一句话总结

本研究提出了一种基于ASR的意大利语学习者英语二语发音错误检测系统,采用基于常见发音错误训练的音素级错误语言模型。通过将错误规则整合到语言模型中,系统将音素错误率(PER)降低至23%,显著提升了成人与儿童学习者基线模型的检测准确率。

ABSTRACT

Pronunciation is one of the fundamentals of language learning, and it is considered a primary factor of spoken language when it comes to an understanding and being understood by others. The persistent presence of high error rates in speech recognition domains resulting from mispronunciations motivates us to find alternative techniques for handling mispronunciations. In this study, we develop a mispronunciation assessment system that checks the pronunciation of non-native English speakers, identifies the commonly mispronounced phonemes of Italian learners of English, and presents an evaluation of the non-native pronunciation observed in phonetically annotated speech corpora. In this work, to detect mispronunciations, we used a phone-based ASR implemented using Kaldi. We used two non-native English labeled corpora; (i) a corpus of Italian adults contains 5,867 utterances from 46 speakers, and (ii) a corpus of Italian children consists of 5,268 utterances from 78 children. Our results show that the selected error model can discriminate correct sounds from incorrect sounds in both native and nonnative speech, and therefore can be used to detect pronunciation errors in non-native speech. The phone error rates show improvement in using the error language model. The ASR system shows better accuracy after applying the error model on our selected corpora.

研究动机与目标

  • 开发一种自动系统,用于在音素层面对非母语英语语音中的发音错误进行检测。
  • 通过整合从意大利语学习者常见发音错误中提取的错误规则,提升ASR在二语学习者中的准确性。
  • 评估错误语言模型在区分成人与儿童学习者正确与错误音素方面的有效性。
  • 通过基于音素转写识别具体发音错误,为学习者提供可操作的反馈。

提出的方法

  • 系统采用基于Kaldi的音素级ASR框架,其声学模型在母语者与非母语者语音数据上进行训练。
  • 通过分析意大利语二语学习者语音中的发音错误,提取替换、删除和插入规则,构建错误语言模型。
  • 将错误规则编码至定制词典中,使ASR系统能够识别并评分非母语发音。
  • 通过比较母语与非母语声学模型的对数似然得分,评估发音质量。
  • 使用音素错误率(PER)作为主要指标,评估不同语言模型配置下的系统性能。
  • 使用两个语料库:46名成人意大利学习者的5,867条语音,以及78名儿童的5,268条语音。

实验结果

研究问题

  • RQ1基于错误优化语言模型的ASR系统能否有效检测二语学习者语音中的发音错误音素?
  • RQ2与标准模型相比,语言模型中包含错误规则对音素错误率(PER)有何影响?
  • RQ3该系统在成人与儿童学习者中,能在多大程度上实现类人水平的音素级错误检测?
  • RQ4意大利语学习者中最常见的发音错误模式是什么,以及如何系统地建模这些模式?

主要发现

  • 错误语言模型实现了最低的音素错误率(PER)23%,在成人与儿童学习者语料中均优于基线模型。
  • 系统在母语与非母语语音中成功区分了正确与错误发音,表现出跨年龄群体的鲁棒性。
  • 在语言模型中引入错误规则显著提升了ASR准确性,尤其在高频错误如/dh/→/d/、/n d/→/n d/和/er/→/er r/方面表现突出。
  • 与先前研究相比,本系统的性能更优,PER为23%,而以往研究在非母语语音数据上的报告WER最高达30%。
  • 通过为如'HER'等词语提供多种转写方式的定制词典,系统更好地识别了非母语变体,提升了反馈准确性。
  • 结果表明,基于ASR的系统可针对特定母语背景(如意大利语)进行有效定制,从而增强计算机辅助发音训练(CAPT)的效果。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。