Skip to main content
QUICK REVIEW

[论文解读] Gender Based Emotion Recognition System for Telugu Rural Dialects Using Hidden Markov Models

P. V. G. D. Prasad Reddy, Ajay Prasad|arXiv (Cornell University)|Jun 23, 2010
Speech Recognition and Synthesis参考文献 9被引用 14
一句话总结

该论文提出了一种基于性别的情绪识别系统,用于特卢古语农村方言,采用隐马尔可夫模型(HMMs),利用39维特征向量(13个MFCCs,13个倒谱系数差分,13个倒谱系数二阶差分)对五种情绪状态——愤怒、惊讶、快乐、悲伤和中性——进行分类。该系统在识别愤怒情绪方面,对两性均实现了80%的准确率,经由人工评估验证,并通过在TRDAP数据库上进行的性别相关与性别无关的实验得到验证。

ABSTRACT

Automatic emotion recognition in speech is a research area with a wide range of applications in human interactions. The basic mathematical tool used for emotion recognition is Pattern recognition which involves three operations, namely, pre-processing, feature extraction and classification. This paper introduces a procedure for emotion recognition using Hidden Markov Models (HMM), which is used to divide five emotional states: anger, surprise, happiness, sadness and neutral state. The approach is based on standard speech recognition technology using hidden continuous markov model by selection of low level features and the design of the recognition system. Emotional Speech Database from Telugu Rural Dialects of Andhra Pradesh (TRDAP) was designed using several speaker's voices comprising the emotional states. The accuracy of recognizing five different emotions for both genders of classification is 80% for anger-emotion which is achieved by using the best combination of 39-dimensioanl feature vector for every frame (13 MFCCs, 13 Delta Coefficients and 13 Acceleration Coefficients) and a classifier using HMM. This outcome very much matches with that acquired with the same database with subjective evaluation by human judges. Both gender-dependent and gender-independent experiments are conducted on TRDAP emotional speech database.

研究动机与目标

  • 开发一种专为特卢古语农村方言设计的自动情绪识别系统,此类方言在语音情绪研究中代表性不足。
  • 解决在资源有限、语言和声学数据匮乏的区域性方言中进行情绪识别的挑战。
  • 评估基于HMM的分类方法在新创建的情感语音数据库上,采用性别相关与性别无关模型的性能表现。
  • 通过连续HMM与低层级声学特征,为印度区域语言的情绪识别建立基准。

提出的方法

  • 该系统采用三阶段模式识别流程:预处理、特征提取和分类。
  • 每帧语音中提取低层级声学特征,包括13个梅尔频率倒谱系数(MFCCs)、13个倒谱系数差分和13个倒谱系数二阶差分,构成39维向量。
  • 使用连续HMM框架,为每种情绪类别训练隐马尔可夫模型(HMMs),以建模语音的时序动态特性。
  • TRDAP数据库包含多位特卢古语农村方言说话人的情感语音,用于模型的训练与测试。
  • 在相同数据集上训练并评估性别相关与性别无关的HMM分类器,以比较性能差异。
  • 系统采用标准语音识别技术并针对情绪分类进行调整,模型选择基于最大似然估计。

实验结果

研究问题

  • RQ1基于HMM的分类方法能否有效识别特卢古语农村方言中的五种不同情绪——愤怒、惊讶、快乐、悲伤和中性?
  • RQ2针对资源有限的区域性语音,性别特定建模如何影响情绪识别的准确率?
  • RQ313维MFCCs、13维倒谱系数差分与13维倒谱系数二阶差分的组合在多大程度上提升了情绪识别性能?
  • RQ4该系统在相同数据集上的表现与人工主观评估相比如何?
  • RQ5在印度区域方言中,情绪识别的最优特征向量与模型配置是什么?

主要发现

  • 所提出的系统在使用最优39维特征向量时,对愤怒情绪的识别准确率达到80%,为五种情绪类别中的最高值。
  • 在TRDAP数据库上,性别相关HMM模型在情绪识别准确率上优于性别无关模型。
  • 系统性能与人工主观评估结果高度接近,验证了其可靠性和鲁棒性。
  • 13个MFCCs、13个倒谱系数差分与13个倒谱系数二阶差分的组合实现了最佳分类性能。
  • 连续HMM框架能有效建模农村特卢古语方言中情绪语音的时序动态特性。
  • TRDAP数据库可作为未来印度区域语言情绪识别研究的有效且可靠的基准。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。