Skip to main content
QUICK REVIEW

[論文レビュー] Predicting speech intelligibility from EEG using a dilated convolutional network.

Bernd Accou, Mohammad Jalilpour-Monesi|arXiv (Cornell University)|May 14, 2021
EEG and Brain-Computer Interfaces参考文献 30被引用数 9
ひとこと要約

本研究では、被験者固有の再トレーニングを必要とせず、EEG信号から話者の理解度を予測する拡張畳み込みニューラルネットワークを導入し、ベースラインモデルを著しく上回る性能を達成した。また、行動的MATRIXテストと強い相関(r=0.59)を示し、未学習の被験者に対するEEGベースの話者受容閾値の客観的で一般化可能な予測を初めて実現した。

ABSTRACT

Objective: Currently, only behavioral speech understanding tests are available, which require active participation of the person. As this is infeasible for certain populations, an objective measure of speech intelligibility is required. Recently, brain imaging data has been used to establish a relationship between stimulus and brain response. Linear models have been successfully linked to speech intelligibility but require per-subject training. We present a deep-learning-based model incorporating dilated convolutions that can be used to predict speech intelligibility without subject-specific (re)training. Methods: We evaluated the performance of the model as a function of input segment length, EEG frequency band and receptive field size while comparing it to a baseline model. Next, we evaluated performance on held-out data and finetuning. Finally, we established a link between the accuracy of our model and the state-of-the-art behavioral MATRIX test. Results: The model significantly outperformed the baseline for every input segment length (p$\leq10^{-9}$), for all EEG frequency bands except the theta band (p$\leq0.001$) and for receptive field sizes larger than 125 ms (p$\leq0.05$). Additionally, finetuning significantly increased the accuracy (p$\leq0.05$) on a held-out dataset. Finally, a significant correlation (r=0.59, p=0.0154) was found between the speech reception threshold estimated using the behavioral MATRIX test and our objective method. Conclusion: Our proposed dilated convolutional model can be used as a proxy for speech intelligibility. Significance: Our method is the first to predict the speech reception threshold from EEG for unseen subjects, contributing to objective measures of speech intelligibility.

研究の動機と目的

  • 行動的テストに参加できない被験者を対象として、EEGを用いた客観的で非侵襲的な話者の理解度評価法の開発。
  • 個々の被験者に特化したトレーニングを必要とする従来の線形モデルの限界を克服し、一般化能力を有する深層学習アプローチを導入すること。
  • 脳活動パターンを用いて信頼性の高い、被験者に依存しない話者受容閾値の代理指標を確立すること。
  • 入力セグメント長、EEG周波数帯、受容 field サイズの変動に対するモデルの頑健性を評価すること。
  • ゴールドスタンダードの行動的MATRIXテストと比較して、モデルの予測精度を検証すること。

提案手法

  • 時間的依存性を捉えるために、拡張された受容 field を有する拡張畳み込みニューラルネットワーク(DCNN)を用い、パrameter数を増加させずに長距離の文脈をモデル化する。
  • 入力構成をセグメント長、周波数帯(例:アルファ、ベータ、ガンマ)および受容 field サイズの変更により変更し、固定長のウィンドウに分割された生または前処理済みEEGデータを処理する。
  • ネットワークが長い時間窓にわたる言語関連神経動態を捉えるために、畳み込みの拡張を用いて有効な受容 field を拡大する。
  • 一般化能力と新規被験者への適応性を評価するため、ホールドアウトテストデータとファインチューニング戦略を用いてモデル性能を評価する。
  • 損失関数を平均二乗誤差または同様の回帰基準により最適化することで、エンドツーエンドで話者の理解度スコアを予測するようにモデルを訓練する。
  • 相関係数などの複数の評価指標を用いて、ベースライン線形モデルと比較して性能をベンチマークする。

実験結果

リサーチクエスチョン

  • RQ1拡張畳み込みを用いた深層学習モデルは、被験者固有の再トレーニングを必要とせず、EEGから話者の理解度を予測できるか?
  • RQ2入力セグメント長は、モデルの話者の理解度予測能力にどのように影響するか?
  • RQ3どのEEG周波数帯が話者の理解度の正確な予測に最も寄与しているか?
  • RQ4言語理解に関連する神経パターンを捉えるために最適な受容 field サイズは何か?
  • RQ5ファインチューニングは、未学習の被験者におけるモデル性能をどの程度向上させるか?

主な発見

  • 拡張畳み込みモデルは、すべての入力セグメント長でベースラインと比べて著しく優れた性能を示した(p ≤ 10⁻⁹)。
  • θ帯を除くすべてのEEG周波数帯で、ベースラインより統計的に有意な性能向上(p ≤ 0.001)を達成した。
  • 125 ms より大きな受容 field サイズが有意に優れた性能を示した(p ≤ 0.05)ことから、長期的な時間的文脈の重要性が示された。
  • ホールドアウトデータ上でファインチューニングを施すことで、モデルの精度が顕著に向上した(p ≤ 0.05)、新規被験者への適応性を示した。
  • モデルが予測した話者受容閾値と行動的MATRIXテストの結果との間に有意な正の相関(r = 0.59, p = 0.0154)が観察された。
  • 本研究は、未学習の被験者に対して、客観的で非侵襲的かつ一般化可能な深層学習アプローチにより、EEGから話者受容閾値を予測する初の手法を提示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。