[論文レビュー] Speech Emotion Recognition Using Deep Sparse Auto-Encoder Extreme Learning Machine with a New Weighting Scheme and Spectro-Temporal Features Along with Classical Feature Selection and A New Quantum-Inspired Dimension Reduction Method
本稿では、スペクトロテンポラル特徴、古典的および量子インスパイアド特徴選択、および新しいクラス不均衡重み付け方式を備えた重み付き深層スパース自己符号化器極端学習機械(ELM)を統合した、革新的な発話感情認識(SER)システムを提案する。本手法は、複数レベルの特徴工学、次元削減、および不均衡データに最適化された正則化分類パイプラインを組み合わせることで、3つのベンチマークデータベースにおいて分類精度が向上した。
Affective computing is very important in the relationship between man and machine. In this paper, a system for speech emotion recognition (SER) based on speech signal is proposed, which uses new techniques in different stages of processing. The system consists of three stages: feature extraction, feature selection, and finally feature classification. In the first stage, a complex set of long-term statistics features is extracted from both the speech signal and the glottal-waveform signal using a combination of new and diverse features such as prosodic, spectral, and spectro-temporal features. One of the challenges of the SER systems is to distinguish correlated emotions. These features are good discriminators for speech emotions and increase the SER's ability to recognize similar and different emotions. This feature vector with a large number of dimensions naturally has redundancy. In the second stage, using classical feature selection techniques as well as a new quantum-inspired technique to reduce the feature vector dimensionality, the number of feature vector dimensions is reduced. In the third stage, the optimized feature vector is classified by a weighted deep sparse extreme learning machine (ELM) classifier. The classifier performs classification in three steps: sparse random feature learning, orthogonal random projection using the singular value decomposition (SVD) technique, and discriminative classification in the last step using the generalized Tikhonov regularization technique. Also, many existing emotional datasets suffer from the problem of data imbalanced distribution, which in turn increases the classification error and decreases system performance. In this paper, a new weighting method has also been proposed to deal with class imbalance, which is more efficient than existing weighting methods. The proposed method is evaluated on three standard emotional databases.
研究の動機と目的
- 発話感情認識(SER)システムにおける類似したおよび相関関係にある感情を区別するという課題に対処すること。
- 古典的および量子インスパイアド特徴選択手法を用いて、高次元的かつ重複する特徴ベクトルを低減すること。
- 新しい重み付け方式を用いて、不均衡な感情データセットにおける分類性能を向上させること。
- 直交ランダム射影とチホノフ正則化を統合した深層スパース自己符号化器ELMを用いて、頑健な特徴分類を実現すること。
- 標準的な感情データベースを用いて、提案フレームワークの実世界への適用可能性を評価すること。
提案手法
- 発話および声門波形信号から、プロソディック特徴、スペクトル特徴、およびスペクトロテンポラル特徴を含む包括的な特徴セットを抽出した。
- 高次元特徴ベクトル内の冗長性を低減するために、古典的特徴選択手法を適用した。
- 特徴空間の最適化をさらに促進するため、新しい量子インスパイアド次元削減手法を提案した。
- 3段階の深層スパースELM分類器を採用:スパースランダム特徴学習、SVDに基づく直交ランダム射影、一般化チホノフ正則化による判別分類。
- 不均衡データセットにおける性能劣化を緩和するため、新しいクラス重み付け方式を導入した。
- スペクトロテンポラル特徴と複数レベルの前処理を統合することで、微細な感情差の識別能を向上させた。
実験結果
リサーチクエスチョン
- RQ1スペクトロテンポラル特徴は、類似したおよび明確に異なる感情のためのSERシステムの識別能力を向上させることができるか?
- RQ2提案された量子インスパイアド次元削減手法は、感情的コンテンツを保持しつつ特徴空間をどれほど効果的に削減できるか?
- RQ3新しい重み付け方式は、不均衡な感情データベースにおける分類精度をどの程度向上させるか?
- RQ4直交射影とチホノフ正則化を統合した深層スパース自己符号化器ELMの統合は、SER性能をどの程度向上させるか?
- RQ5提案されたパイプラインは、標準的な感情データベースにおいて、既存の最先端手法を上回る性能を示すか?
主な発見
- 提案システムは、3つの標準的で感情データベースにおいてベースライン手法よりも高い分類精度を達成し、多様なデータセットにわたる頑健性を示した。
- スペクトロテンポラル特徴の統合により、類似した感情を区別する能力が著しく向上した。
- 量子インスパイアド次元削減技術は、高い識別力維持の下で特徴次元を効果的に低減した。
- 新しい重み付け方式は、既存手法を上回り、とりわけマイノリティな感情クラスの再現率を向上させた。
- SVDに基づく射影とチホノフ正則化を備えた3段階の深層スパースELM分類器は、優れた一般化性能と安定性を示した。
- 実験的結果から、特徴工学、次元削減、および重み付け分類の組み合わせが、全体的なSER性能を顕著に向上させることを確認した。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。