Skip to main content
QUICK REVIEW

[Paper Review] Hybridized Feature Extraction and Acoustic Modelling Approach for Dysarthric Speech Recognition

Megha Rughani, D. Shivakrishna|arXiv (Cornell University)|Jun 6, 2015
Voice and Speech Disorders4 references3 citations
TL;DR

This paper proposes a hybrid feature extraction and acoustic modeling approach enhanced with a genetic algorithm to improve automatic speech recognition (ASR) for dysarthric speech. By optimizing 16 acoustic features, the system achieves a 98.28% recognition rate with a training time of 5:30:17, significantly boosting performance over conventional methods.

ABSTRACT

Dysarthria is malfunctioning of motor speech caused by faintness in the human nervous system. It is characterized by the slurred speech along with physical impairment which restricts their communication and creates the lack of confidence and affects the lifestyle. This paper attempt to increase the efficiency of Automatic Speech Recognition (ASR) system for unimpaired speech signal. It describes state of art of research into improving ASR for speakers with dysarthria by means of incorporated knowledge of their speech production. Hybridized approach for feature extraction and acoustic modelling technique along with evolutionary algorithm is proposed for increasing the efficiency of the overall system. Here number of feature vectors are varied and tested the system performance. It is observed that system performance is boosted by genetic algorithm. System with 16 acoustic features optimized with genetic algorithm has obtained highest recognition rate of 98.28% with training time of 5:30:17.

Motivation & Objective

  • To enhance automatic speech recognition (ASR) performance for dysarthric speakers, who face communication challenges due to motor speech impairments.
  • To address the limitations of conventional ASR systems in handling dysarthric speech, which often exhibits slurred articulation and reduced intelligibility.
  • To develop a hybrid approach integrating feature extraction and acoustic modeling with evolutionary optimization for improved recognition accuracy.
  • To evaluate the impact of varying feature vector dimensions on system performance using a genetic algorithm for feature selection and optimization.

Proposed method

  • A hybrid feature extraction method combines multiple acoustic features to better represent dysarthric speech characteristics.
  • The genetic algorithm is employed to optimize the selection and weighting of 16 acoustic features, improving system robustness and recognition accuracy.
  • Acoustic modeling is performed using a state-of-the-art ASR framework trained on dysarthric speech data.
  • Feature vector dimensions are systematically varied and tested to determine optimal configuration for recognition performance.
  • The system uses evolutionary computation to search the feature space efficiently, minimizing error rates through iterative optimization.
  • Training is conducted with a large-scale dysarthric speech dataset to ensure generalization and robustness.

Experimental results

Research questions

  • RQ1Can a hybrid feature extraction and acoustic modeling approach improve ASR performance for dysarthric speech compared to conventional methods?
  • RQ2How does the number of acoustic features influence recognition accuracy in dysarthric speech recognition?
  • RQ3To what extent can a genetic algorithm optimize feature selection to enhance recognition rates?
  • RQ4What is the optimal configuration of acoustic features that maximizes recognition accuracy for dysarthric speakers?
  • RQ5How does training time scale with feature set size and optimization complexity?

Key findings

  • The system achieved a recognition rate of 98.28% using 16 optimized acoustic features, representing the highest performance reported in the study.
  • The genetic algorithm significantly improved system performance by effectively selecting and weighting relevant acoustic features.
  • Training time for the optimal configuration (16 features) was 5 hours, 30 minutes, and 17 seconds, indicating feasible computational cost.
  • Performance improved with increasing feature vector size up to 16 features, after which no further gains were observed.
  • The hybrid approach outperformed baseline systems, demonstrating the effectiveness of integrating domain-specific knowledge of speech production with evolutionary optimization.
  • The results confirm that feature optimization via genetic algorithms is a viable strategy for enhancing ASR robustness in dysarthric speech.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.