Skip to main content
QUICK REVIEW

[論文レビュー] Voiceprint recognition of Parkinson patients based on deep learning

Zhijing Xu, Juan Wang|arXiv (Cornell University)|Jan 1, 2018
Voice and Speech Disorders被引用数 1
ひとこと要約

本研究では、声紋特徴抽出に重み付きメル周波数ケプストラム係数(WMFCC)を組み合わせ、パーキンソン病(PD)患者と健常者を分類するための深層ニューラルネットワーク(DNN)とミニバッチ勾配降下法(MBGD)を用いた深層学習手法を提案する。持続音声記録を用いた実験で、89.5%の精度を達成し、SVMのような従来手法を上回る性能を示した。

ABSTRACT

More than 90% of the Parkinson Disease (PD) patients suffer from vocal disorders. Speech impairment is already indicator of PD. This study focuses on PD diagnosis through voiceprint features. In this paper, a method based on Deep Neural Network (DNN) recognition and classification combined with Mini-Batch Gradient Descent (MBGD) is proposed to distinguish PD patients from healthy people using voiceprint features. In order to exact the voiceprint features from patients, Weighted Mel Frequency Cepstrum Coefficients (WMFCC) is applied. The proposed method is tested on experimental data obtained by the voice recordings of three sustained vowels /a/, /o/ and /u/ from participants (48 PD and 20 healthy people). The results show that the proposed method achieves a high accuracy of diagnosis of PD patients from healthy people, than the conventional methods like Support Vector Machine (SVM) and other mentioned in this paper. The accuracy achieved is 89.5%. WMFCC approach can solve the problem that the high-order cepstrum coefficients are small and the features component's representation ability to the audio is weak. MBGD reduces the computational loads of the loss function, and increases the training speed of the system. DNN classifier enhances the classification ability of voiceprint features. Therefore, the above approaches can provide a solid solution for the quick auxiliary diagnosis of PD in early stage.

研究の動機と目的

  • 非侵襲的声ベースのバイオマーカーを用いた、早期パーキンソン病(PD)診断の課題に取り組む。
  • 声紋特徴を用いて、PD患者と健常者を分類する精度を向上させる。
  • 高次ケプストラム係数の表現が弱いという従来の特徴抽出手法の限界を克服する。
  • 深層学習フレームワーク内でのミニバッチ勾配降下法により、計算負荷を低減し、学習を高速化する。

提案手法

  • 持続音声 /a/、/o/、/u/ から、重み付きメル周波数ケプストラム係数(WMFCC)を用いて声紋特徴を抽出する。
  • WMFCC手法は、高次ケプストラム係数が小さい問題を解消することで、特徴表現を向上させる。
  • 分類性能を向上させるために、ミニバッチ勾配降下法(MBGD)を用いて深層ニューラルネットワーク(DNN)分類器を学習する。
  • MBGDは、DNN学習中の計算負荷を低減し、収束を高速化する。
  • 48例のPD患者と20名の健常者から成るデータセットを、持続音声記録を用いて学習およびテストした。
  • DNNモデルは、PDと健常対照群を区別するための声紋特徴の階層的表現を学習する。

実験結果

リサーチクエスチョン

  • RQ1従来のMFCCと比較して、WMFCCはパーキンソン病検出のための声紋特徴の識別力を向上させることができるか?
  • RQ2SVMのような従来の分類器と比較して、MBGDを用いたDNNは、声サンプルからのPD診断においてどのように性能を発揮するか?
  • RQ3提案手法は、高い診断精度を維持しつつ、計算負荷をどの程度低減できるか?
  • RQ4持続音声発話は、深層学習を用いた早期パーキンソン病診断に信頼性を持って利用可能か?

主な発見

  • 提案手法は、PD患者と健常者を区別する診断精度が89.5%に達した。
  • WMFCCの使用により、特に高次ケプストラム係数の特徴表現が顕著に向上した。
  • ミニバッチ勾配降下法は、DNNモデルの計算負荷を低減し、学習速度を向上させた。
  • DNN分類器は、SVMのような従来手法よりも優れた分類性能を示した。
  • 本システムは、信頼性のある早期パーキンソン病検出のため、持続音声記録を効果的に活用した。
  • WMFCCとDNNをMBGDで統合することで、自動PDスクリーニングに耐性があり、効率的なソリューションが得られた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。