[論文レビュー] Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?
本研究では、オープンソースのKaldi-NLを用いて、音声認識をリアルタイムで実行する自動音声聴覚検査システムを提案した。30名の聴力正常とされるオランダ語話者を対象としたデジットインノイズ(DIN)テストにおいて、平均語誤り率(WER)は5.0%であった。シミュレーションの結果、最大4つの誤ったトリプレットが存在しても、SRTの変動は通常の被験者内変動(0.70 dB)の範囲内に収まり、人為的監視なしで聴力スクリーニングに臨床的に実用可能であることが示された。
A practical speech audiometry tool is the digits-in-noise (DIN) test for hearing screening of populations of varying ages and hearing status. The test is usually conducted by a human supervisor (e.g., clinician), who scores the responses spoken by the listener, or online, where software scores the responses entered by the listener. The test has 24-digit triplets presented in an adaptive staircase procedure, resulting in a speech reception threshold (SRT). We propose an alternative automated DIN test setup that can evaluate spoken responses whilst conducted without a human supervisor, using the open-source automatic speech recognition toolkit, Kaldi-NL. Thirty self-reported normal-hearing Dutch adults (19-64 years) completed one DIN + Kaldi-NL test. Their spoken responses were recorded and used for evaluating the transcript of decoded responses by Kaldi-NL. Study 1 evaluated the Kaldi-NL performance through its word error rate (WER), percentage of summed decoding errors regarding only digits found in the transcript compared to the total number of digits present in the spoken responses. Average WER across participants was 5.0% (range 0-48%, SD = 8.8%), with average decoding errors in three triplets per participant. Study 2 analyzed the effect that triplets with decoding errors from Kaldi-NL had on the DIN test output (SRT), using bootstrapping simulations. Previous research indicated 0.70 dB as the typical within-subject SRT variability for normal-hearing adults. Study 2 showed that up to four triplets with decoding errors produce SRT variations within this range, suggesting that our proposed setup could be feasible for clinical applications.
研究の動機と目的
- オープンソースのASRを用いて、低コストで自動化された、従来の人為的監視を要する音声聴覚検査の代替手段を開発すること。
- 事前学習済みのKaldi-NLの性能を、聴力スクリーニングの目的でノイズ環境下でのデジットトリプレット認識に評価すること。
- ASRのデコード誤りがDINテストにおける最終的な音声受容閾値(SRT)に与える影響を評価すること。
- 認識誤りが存在する状況下でも、Kaldi-NLを用いた自動SRT推定が臨床的に信頼できるかどうかを検証すること。
提案手法
- 30名の自己申告による聴力正常なオランダ語話者に、録音された反応を用いてDINテストを実施し、ASR評価用に使用した。
- オープンソースのKaldi-NLを用いて、録音されたデジットトリプレットの音声をデコードし、特にデジット認識の正確性に注目した。
- 参加者ごとに語誤り率(WER)を算出し、ターゲットとなる音声項目に特化した、デジットのみの誤りを測定することで、性能を隔離して評価した。
- ブートストラップ法を用いてSRT推定をシミュレートし、最大4つのトリプレットにデコード誤りが含まれる場合の変動を評価した。
- 聴力正常な被験者における既知の被験者内SRT変動(0.70 dB)と比較して、誤り状況下でのSRT変動を評価した。
- 実際の誤り分布を想定した統計的シミュレーションを用いて、自動SRT出力の頑健性を評価した。
実験結果
リサーチクエスチョン
- RQ1オープンソースのKaldi-NLは、DINテストにおける信頼できるデジット認識のための十分に低い語誤り率を達成できるか?
- RQ2デジットトリプレットにおけるデコード誤りは、自動聴覚検査における最終的な音声受容閾値(SRT)にどのように影響するか?
- RQ3ASR誤りによるSRTの変動は、自然な被験者内変動と比較してどの程度の影響を及ぼすか?
- RQ4Kaldi-NLを基盤とする自動テストによるSRT出力は、臨床的用途に耐えうるほど安定しているか?
主な発見
- 参加者間での平均語誤り率(WER)は5.0%であり、標準偏差は8.8%、範囲は0%から48%であった。
- 平均して、各参加者に3つのデジットトリプレットにデコード誤りが生じており、中程度の誤り率であり、管理可能であることが示された。
- シミュレーションの結果、最大4つの誤ったトリプレットが存在しても、SRTの変動は通常の被験者内SRT変動(0.70 dB)の範囲内に収まった。
- 結果として、認識誤りが存在する状況下でも、Kaldi-NLを用いた自動DIN検査は臨床的に信頼できるSRT推定を可能にする。
- 本システムは、人為的監視なしに、リソースが限られた地域や遠隔地での聴力スクリーニングへの導入が可能であることが示された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。