Skip to main content
QUICK REVIEW

[논문 리뷰] Speech Emotion Recognition Using Deep Sparse Auto-Encoder Extreme Learning Machine with a New Weighting Scheme and Spectro-Temporal Features Along with Classical Feature Selection and A New Quantum-Inspired Dimension Reduction Method

Fatemeh Daneshfar, Seyed Jahanshah Kabudian|arXiv (Cornell University)|2021. 11. 13.
Machine Learning and ELM인용 수 6
한 줄 요약

이 논문은 음성 정서 인식(SER) 시스템에 스펙트로템포럴 특징, 고전적 및 양자 기반 특징 선택, 그리고 새로운 클래스 불균형 가중치 기반 가중치가 부여된 딥 스퍼스 오토에코더 극단적 학습기계(ELM)를 통합한 새로운 방법을 제안한다. 이 방법은 다중 수준의 특징 공학, 차원 축소, 그리고 불균형 데이터에 최적화된 정규화된 분류 파ipeline을 통해 세 가지 기준 데이터베이스에서 분류 정확도를 향상시킨다.

ABSTRACT

Affective computing is very important in the relationship between man and machine. In this paper, a system for speech emotion recognition (SER) based on speech signal is proposed, which uses new techniques in different stages of processing. The system consists of three stages: feature extraction, feature selection, and finally feature classification. In the first stage, a complex set of long-term statistics features is extracted from both the speech signal and the glottal-waveform signal using a combination of new and diverse features such as prosodic, spectral, and spectro-temporal features. One of the challenges of the SER systems is to distinguish correlated emotions. These features are good discriminators for speech emotions and increase the SER's ability to recognize similar and different emotions. This feature vector with a large number of dimensions naturally has redundancy. In the second stage, using classical feature selection techniques as well as a new quantum-inspired technique to reduce the feature vector dimensionality, the number of feature vector dimensions is reduced. In the third stage, the optimized feature vector is classified by a weighted deep sparse extreme learning machine (ELM) classifier. The classifier performs classification in three steps: sparse random feature learning, orthogonal random projection using the singular value decomposition (SVD) technique, and discriminative classification in the last step using the generalized Tikhonov regularization technique. Also, many existing emotional datasets suffer from the problem of data imbalanced distribution, which in turn increases the classification error and decreases system performance. In this paper, a new weighting method has also been proposed to deal with class imbalance, which is more efficient than existing weighting methods. The proposed method is evaluated on three standard emotional databases.

연구 동기 및 목표

  • 유사하고 상관관계가 높은 정서를 구분하는 데 어려움을 겪는 음성 정서 인식(SER) 시스템의 문제를 해결하기 위해.
  • 고전적 및 양자 기반 특징 선택 기법을 통해 고차원적이고 중복된 특징 벡터를 줄이기 위해.
  • 새로운 가중치 기반 기법을 사용하여 불균형 정서 데이터셋에서의 분류 성능을 향상시키기 위해.
  • 직교 랜덤 프로젝션과 티코노프 정규화를 통합한 딥 스퍼스 오토에코더 ELM을 통해 강력한 특징 분류 성능을 확보하기 위해.
  • 실제 응용 가능성을 평가하기 위해 표준 정서 데이터베이스에서 제안된 프레임워크를 평가하기 위해.

제안 방법

  • 음성 신호와 글로탈 웨이브폼 신호에서 모두 주로드릭, 스펙트럴, 스펙트로템포럴 특징을 포함한 포괄적인 특징 세트를 추출하였다.
  • 고차원 특징 벡터의 중복성을 줄이기 위해 고전적 특징 선택 기법을 적용하였다.
  • 특징 공간을 추가로 최적화하기 위해 새로운 양자 기반 차원 축소 방법을 제안하였다.
  • 세 단계로 구성된 딥 스퍼스 ELM 분류기: 스퍼스 랜덤 특징 학습, SVD 기반 직교 랜덤 프로젝션, 일반화된 티코노프 정규화를 통한 분류.
  • 불균형 데이터셋에서 성능 저하를 완화하기 위해 새로운 클래스 가중치 기반 기법을 도입하였다.
  • 스펙트로템포럴 특징과 다중 수준 전처리를 통합하여 미묘한 정서적 차이를 구분하는 능력을 향상시켰다.

실험 결과

연구 질문

  • RQ1스펙트로템포럴 특징은 비슷하거나 다릅니까만 정서를 구분하는 데 SER 시스템의 구분 능력을 향상시키는가?
  • RQ2제안된 양자 기반 차원 축소 방법은 정서적 내용을 유지하면서 특징 공간을 얼마나 효과적으로 줄이는가?
  • RQ3새로운 가중치 기반 기법은 불균형 정서 데이터베이스에서 분류 정확도를 얼마나 향상시키는가?
  • RQ4직교 프로젝션과 티코노프 정규화를 통합한 딥 스퍼스 오토에코더 ELM의 통합은 어떻게 SER 성능을 향상시키는가?
  • RQ5제안된 파이프라인은 표준 정서 데이터베이스에서 기존 최고 수준의 방법들을 초월하는가?

주요 결과

  • 제안된 시스템은 세 가지 표준 정서 데이터베이스에서 기준 방법보다 높은 분류 정확도를 달성하여 다양한 데이터셋에서의 강건성을 입증하였다.
  • 스펙트로템포럴 특징의 통합은 비슷한 정서를 구분하는 능력을 시스템이 크게 향상시켰다.
  • 양자 기반 차원 축소 기법은 높은 구분 능력을 유지하면서 특징 차원을 효과적으로 줄였다.
  • 새로운 가중치 기반 기법은 기존 방법보다 더 나은 성능을 보였으며, 특히 소수 정서 클래스의 재현율을 향상시켰다.
  • SVD 기반 프로젝션과 티코노프 정규화를 통합한 세 단계 딥 스퍼스 ELM 분류기는 뛰어난 일반화 능력과 안정성을 제공하였다.
  • 실증 결과는 특징 공학, 차원 축소, 가중치 기반 분류의 조합이 전반적인 SER 성능을 크게 향상시킨다는 것을 확인하였다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.