[论文解读] Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?
本研究提出了一种使用开源 Kaldi-NL 的自动语音测听系统,用于在噪声中识别数字(DIN)测试中的实时语音识别,30 名听力正常荷兰成人受试者的平均词错误率(WER)为 5.0%。模拟结果显示,最多四个错误的三重组导致的 SRT 变化在典型个体内变异性范围(0.70 dB)内,表明该系统在无需人工监督的情况下具备临床可行性,可用于听力筛查。
A practical speech audiometry tool is the digits-in-noise (DIN) test for hearing screening of populations of varying ages and hearing status. The test is usually conducted by a human supervisor (e.g., clinician), who scores the responses spoken by the listener, or online, where software scores the responses entered by the listener. The test has 24-digit triplets presented in an adaptive staircase procedure, resulting in a speech reception threshold (SRT). We propose an alternative automated DIN test setup that can evaluate spoken responses whilst conducted without a human supervisor, using the open-source automatic speech recognition toolkit, Kaldi-NL. Thirty self-reported normal-hearing Dutch adults (19-64 years) completed one DIN + Kaldi-NL test. Their spoken responses were recorded and used for evaluating the transcript of decoded responses by Kaldi-NL. Study 1 evaluated the Kaldi-NL performance through its word error rate (WER), percentage of summed decoding errors regarding only digits found in the transcript compared to the total number of digits present in the spoken responses. Average WER across participants was 5.0% (range 0-48%, SD = 8.8%), with average decoding errors in three triplets per participant. Study 2 analyzed the effect that triplets with decoding errors from Kaldi-NL had on the DIN test output (SRT), using bootstrapping simulations. Previous research indicated 0.70 dB as the typical within-subject SRT variability for normal-hearing adults. Study 2 showed that up to four triplets with decoding errors produce SRT variations within this range, suggesting that our proposed setup could be feasible for clinical applications.
研究动机与目标
- 开发一种低成本、自动化的语音测听替代方案,取代传统的人工监督语音测听,采用开源自动语音识别(ASR)技术。
- 评估预训练的 Kaldi-NL 在噪声条件下识别数字三重组的性能,以用于听力筛查。
- 评估 ASR 解码错误对 DIN 测试中最终语音接收阈值(SRT)的影响。
- 确定在存在识别错误的情况下,基于 Kaldi-NL 的自动 SRT 估计是否仍具有临床可靠性。
提出的方法
- 使用录制的语音响应对 30 名自我报告听力正常的荷兰成人受试者实施 DIN 测试,用于 ASR 评估。
- 使用开源 Kaldi-NL 对录制语音中的数字三重组进行解码,重点评估数字识别的准确性。
- 计算每位受试者的词错误率(WER),仅测量数字部分的错误,以隔离目标语音项目的性能表现。
- 通过 resampling 方法模拟 SRT 估计,评估当最多四个三重组包含解码错误时的变异性。
- 将误差条件下的 SRT 变异性与听力正常成人已知的个体内 SRT 变异性(0.70 dB)进行比较。
- 使用统计模拟方法评估在现实误差分布下,自动 SRT 输出的鲁棒性。
实验结果
研究问题
- RQ1开源 Kaldi-NL 是否能在 DIN 测试中实现足够低的词错误率,以确保数字识别的可靠性?
- RQ2数字三重组中的解码错误如何影响自动测听中的最终语音接收阈值(SRT)?
- RQ3与自然的个体内变异性相比,ASR 错误在多大程度上影响 SRT 的变异性?
- RQ4基于 Kaldi-NL 的自动测试所输出的 SRT 是否足够稳定,可应用于临床场景?
主要发现
- 受试者之间的平均词错误率(WER)为 5.0%,标准差为 8.8%,范围为 0% 至 48%。
- 平均每位受试者有三个数字三重组出现解码错误,表明错误率中等但可管理。
- 模拟结果表明,最多四个错误三重组导致的 SRT 变化在典型个体内 SRT 变异性范围(0.70 dB)内。
- 结果表明,尽管存在识别错误,基于 Kaldi-NL 的自动 DIN 测试仍能产生具有临床可靠性的 SRT 估计值。
- 该系统在无需人工监督的低资源或偏远听力筛查环境中具备部署可行性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。