[论文解读] Voiceprint recognition of Parkinson patients based on deep learning
本研究提出一种深度学习方法,结合加权梅尔频率倒谱系数(WMFCC)进行语音特征提取,并采用带有小批量梯度下降(MBGD)的深度神经网络(DNN)对帕金森病(PD)患者与健康个体进行分类。该方法在使用持续元音录音时,区分PD患者的准确率达到89.5%,优于传统的SVM等方法。
More than 90% of the Parkinson Disease (PD) patients suffer from vocal disorders. Speech impairment is already indicator of PD. This study focuses on PD diagnosis through voiceprint features. In this paper, a method based on Deep Neural Network (DNN) recognition and classification combined with Mini-Batch Gradient Descent (MBGD) is proposed to distinguish PD patients from healthy people using voiceprint features. In order to exact the voiceprint features from patients, Weighted Mel Frequency Cepstrum Coefficients (WMFCC) is applied. The proposed method is tested on experimental data obtained by the voice recordings of three sustained vowels /a/, /o/ and /u/ from participants (48 PD and 20 healthy people). The results show that the proposed method achieves a high accuracy of diagnosis of PD patients from healthy people, than the conventional methods like Support Vector Machine (SVM) and other mentioned in this paper. The accuracy achieved is 89.5%. WMFCC approach can solve the problem that the high-order cepstrum coefficients are small and the features component's representation ability to the audio is weak. MBGD reduces the computational loads of the loss function, and increases the training speed of the system. DNN classifier enhances the classification ability of voiceprint features. Therefore, the above approaches can provide a solid solution for the quick auxiliary diagnosis of PD in early stage.
研究动机与目标
- 为解决利用非侵入性语音生物标志物进行早期帕金森病(PD)诊断的挑战。
- 通过语音特征提升PD患者与健康个体分类的准确性。
- 克服传统特征提取方法的局限性,例如高阶倒谱系数表示能力弱的问题。
- 通过深度学习框架中的小批量梯度下降(MBGD)降低计算负载并加速训练。
提出的方法
- 使用加权梅尔频率倒谱系数(WMFCC)从持续元音/a/、/o/和/u/中提取语音特征。
- WMFCC方法通过解决高阶倒谱系数数值过小的问题,增强了特征表示能力。
- 采用小批量梯度下降(MBGD)训练深度神经网络(DNN)分类器,以提升分类性能。
- MBGD降低了计算负载,并加速了DNN训练过程中的收敛速度。
- 系统在包含48名PD患者和20名健康个体的语音数据集上,使用持续元音录音进行训练与测试。
- DNN模型学习语音特征的分层表示,以有效区分PD患者与健康对照组。
实验结果
研究问题
- RQ1与传统的MFCC相比,WMFCC是否能提升语音特征在帕金森病检测中的判别能力?
- RQ2采用MBGD的DNN与传统分类器(如SVM)相比,在基于语音样本的PD诊断中表现如何?
- RQ3所提出的方法在保持高诊断准确率的同时,能在多大程度上降低计算负载?
- RQ4持续元音发音是否能可靠支持使用深度学习进行早期帕金森病诊断?
主要发现
- 所提方法在区分帕金森病患者与健康个体时,诊断准确率达到89.5%。
- 使用WMFCC显著改善了特征表示,尤其在高阶倒谱系数方面表现更优。
- 小批量梯度下降(MBGD)降低了计算负载,并提升了DNN模型的训练速度。
- DNN分类器相比传统方法(如SVM)展现出更优的分类性能。
- 该系统有效利用持续元音录音,实现可靠且早期的帕金森病检测。
- WMFCC与DNN结合MBGD,为自动化帕金森病筛查提供了鲁棒且高效的解决方案。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。