[论文解读] Speech Emotion Recognition Using Deep Sparse Auto-Encoder Extreme Learning Machine with a New Weighting Scheme and Spectro-Temporal Features Along with Classical Feature Selection and A New Quantum-Inspired Dimension Reduction Method
该论文提出了一种新颖的语音情感识别(SER)系统,整合了时频特征、经典与量子启发式特征选择,以及带新型类别不平衡加权方案的加权深度稀疏自编码器极限学习机(ELM)。通过结合多层次特征工程、降维处理和针对不平衡数据优化的正则化分类流程,该方法在三个基准数据库上实现了更高的分类准确率。
Affective computing is very important in the relationship between man and machine. In this paper, a system for speech emotion recognition (SER) based on speech signal is proposed, which uses new techniques in different stages of processing. The system consists of three stages: feature extraction, feature selection, and finally feature classification. In the first stage, a complex set of long-term statistics features is extracted from both the speech signal and the glottal-waveform signal using a combination of new and diverse features such as prosodic, spectral, and spectro-temporal features. One of the challenges of the SER systems is to distinguish correlated emotions. These features are good discriminators for speech emotions and increase the SER's ability to recognize similar and different emotions. This feature vector with a large number of dimensions naturally has redundancy. In the second stage, using classical feature selection techniques as well as a new quantum-inspired technique to reduce the feature vector dimensionality, the number of feature vector dimensions is reduced. In the third stage, the optimized feature vector is classified by a weighted deep sparse extreme learning machine (ELM) classifier. The classifier performs classification in three steps: sparse random feature learning, orthogonal random projection using the singular value decomposition (SVD) technique, and discriminative classification in the last step using the generalized Tikhonov regularization technique. Also, many existing emotional datasets suffer from the problem of data imbalanced distribution, which in turn increases the classification error and decreases system performance. In this paper, a new weighting method has also been proposed to deal with class imbalance, which is more efficient than existing weighting methods. The proposed method is evaluated on three standard emotional databases.
研究动机与目标
- 为解决语音情感识别(SER)系统中难以区分相似且相关情感的挑战。
- 通过经典与量子启发式特征选择技术,降低高维冗余特征向量的维度。
- 通过新型加权方案提升在类别不平衡情感数据集上的分类性能。
- 将深度稀疏自编码器ELM与正交随机投影及Tikhonov正则化相结合,实现鲁棒的特征分类。
- 在标准情感数据库上评估所提出的框架,以验证其在真实场景中的适用性。
提出的方法
- 从语音信号和声门波形信号中提取了包括语音特征、谱特征及时频特征在内的综合性特征集。
- 应用经典特征选择技术以减少高维特征向量中的冗余性。
- 提出一种新型量子启发式降维方法,以进一步优化特征空间。
- 采用三阶段深度稀疏ELM分类器:稀疏随机特征学习、基于SVD的正交随机投影,以及通过广义Tikhonov正则化实现的判别性分类。
- 引入一种新型类别加权方案,以缓解不平衡数据集中性能下降的问题。
- 将时频特征与多级预处理相结合,以增强对细微情感差异的判别能力。
实验结果
研究问题
- RQ1时频特征能否提升SER系统对相似与不同情感的判别能力?
- RQ2所提出的量子启发式降维方法在降低特征空间维度的同时,对保留情感内容的有效性如何?
- RQ3新型加权方案在处理类别不平衡的情感数据库时,对提升分类准确率的改善程度如何?
- RQ4将深度稀疏自编码器ELM与正交投影及Tikhonov正则化相结合,对提升SER性能的贡献有多大?
- RQ5所提出的流程是否在标准情感数据库上优于现有的最先进方法?
主要发现
- 所提系统在三个标准情感数据库上的分类准确率高于基线方法,证明了其在多样化数据集上的鲁棒性。
- 时频特征的整合显著增强了系统区分相似情感的能力。
- 量子启发式降维技术在有效降低特征维度的同时,保持了较高的判别能力。
- 新型加权方案在处理类别不平衡方面优于现有方法,尤其显著提升了少数类情感的召回率。
- 结合SVD投影与Tikhonov正则化的三阶段深度稀疏ELM分类器展现出更优的泛化能力与稳定性。
- 实证结果证实,特征工程、降维处理与加权分类的结合显著提升了整体SER性能。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。