[Paper Review] Assessment of Parkinson's Disease Medication State through Automatic Speech Analysis
This study proposes a speaker-dependent deep learning system that classifies Parkinson’s disease medication states (ON/OFF) using automatic speech analysis. By leveraging acoustic-prosodic features from speech recordings—particularly eGeMAPS and MFCCs—it achieves 95.27% accuracy in semi-spontaneous storytelling tasks, demonstrating strong potential for remote, daily monitoring of PD patients' medication states.
Parkinson's disease (PD) is a progressive degenerative disorder of the central nervous system characterized by motor and non-motor symptoms. As the disease progresses, patients alternate periods in which motor symptoms are mitigated due to medication intake (ON state) and periods with motor complications (OFF state). The time that patients spend in the OFF condition is currently the main parameter employed to assess pharmacological interventions and to evaluate the efficacy of different active principles. In this work, we present a system that combines automatic speech processing and deep learning techniques to classify the medication state of PD patients by leveraging personal speech-based bio-markers. We devise a speaker-dependent approach and investigate the relevance of different acoustic-prosodic feature sets. Results show an accuracy of 90.54% in a test task with mixed speech and an accuracy of 95.27% in a semi-spontaneous speech task. Overall, the experimental assessment shows the potentials of this approach towards the development of reliable, remote daily monitoring and scheduling of medication intake of PD patients.
Motivation & Objective
- To investigate whether speech-based bio-markers can reliably distinguish between ON and OFF medication states in Parkinson’s disease patients.
- To develop a speaker-dependent system that leverages automatic speech processing and deep learning for remote monitoring of PD patients.
- To evaluate the impact of different speech tasks (e.g., reading, storytelling) on the accuracy of medication state classification.
- To assess the relevance of various acoustic-prosodic feature sets (MFCC, delta-MFCC, eGeMAPS) in classifying medication states.
- To explore the feasibility of using limited patient speech data to train accurate, personalized models for clinical monitoring.
Proposed method
- A speaker-dependent deep neural network (DNN) model is trained to classify medication states using personal speech data from each patient.
- Multiple acoustic-prosodic feature sets are extracted: MFCCs, delta-MFCCs, and eGeMAPS, representing spectral, dynamic, and prosodic characteristics of speech.
- Principal component analysis (PCA) is applied to reduce dimensionality while preserving 95% of the variance in the original feature space.
- The system is evaluated on three speech tasks: phonation of /a/, reading a short text, and guided semi-spontaneous storytelling.
- Model performance is measured using utterance-level accuracy across different feature sets and speech tasks.
- The approach uses the FraLusoPark corpus, which includes 74 Portuguese PD patients recorded in both drug-naive and medicated states.
Experimental results
Research questions
- RQ1Can automatic speech analysis reliably classify Parkinson’s disease patients’ medication states (ON vs. OFF) using personal speech-based bio-markers?
- RQ2How does the choice of speech task (e.g., reading vs. storytelling) affect the accuracy of medication state classification?
- RQ3Which acoustic-prosodic feature set (MFCC, delta-MFCC, eGeMAPS) yields the highest classification accuracy for medication state detection?
- RQ4To what extent does dimensionality reduction via PCA impact model performance and efficiency?
- RQ5Can a speaker-dependent model achieve high accuracy with limited patient speech data, enabling practical remote monitoring?
Key findings
- The system achieves 90.54% accuracy in classifying medication states using mixed speech tasks, demonstrating strong baseline performance.
- The highest accuracy of 95.27% is achieved on the semi-spontaneous storytelling task using eGeMAPS features, highlighting the value of natural speech for classification.
- MFCC-based models with PCA reduction achieve 86.49–94.59% accuracy across tasks, with the best performance on storytelling.
- The use of PCA reduces feature dimensionality by up to 79.7% but leads to lower performance for eGeMAPS and MFCC+delta systems, suggesting trade-offs in compression.
- Speaker-dependent models show high per-speaker accuracy (mean 83.78–95.27%), with standard deviation between 12.51% and 14.67%, indicating consistent performance across individuals.
- The results confirm that semi-spontaneous, natural speech recordings significantly outperform controlled tasks like vowel phonation in detecting medication state changes.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.