Skip to main content
QUICK REVIEW

[论文解读] Affect Intensity Estimation Using Multiple Modalities

Amol Patwardhan, Gerald M. Knapp|arXiv (Cornell University)|Jul 5, 2016
Emotion and Mood Recognition被引用 3
一句话总结

本文提出了一种多模态情感强度估计模型,通过加权求和分类置信度、特征点位移和运动速度,融合面部、身体、手势和语音线索。结果表明,在基于0–1唤醒度量表上,语音和手势模态显著提升了情感强度估计的准确性,优于单模态方法。

ABSTRACT

One of the challenges in affect recognition is accurate estimation of the emotion intensity level. This research proposes development of an affect intensity estimation model based on a weighted sum of classification confidence levels, displacement of feature points and speed of feature point motion. The parameters of the model were calculated from data captured using multiple modalities such as face, body posture, hand movement and speech. A preliminary study was conducted to compare the accuracy of the model with the annotated intensity levels. An emotion intensity scale ranging from 0 to 1 along the arousal dimension in the emotion space was used. Results indicated speech and hand modality significantly contributed in improving accuracy in emotion intensity estimation using the proposed model.

研究动机与目标

  • 为解决在人机交互中准确估计情感强度的挑战。
  • 开发一种整合多种模态(面部、身体、手势和语音)以提升情感强度估计的模型。
  • 评估各模态对整体情感强度预测准确率的贡献。
  • 利用Kinect和音频传感器捕获的数据校准模型参数。
  • 在连续的0–1唤醒度量表上,以人工标注的情感强度水平验证模型。

提出的方法

  • 该模型将情感强度计算为各模态分类置信度水平的加权和。
  • 通过基于Kinect的追踪方法,从面部和身体运动中提取特征点位移和运动速度。
  • 对语音特征进行处理,以估计情感强度,进而贡献于最终的加权得分。
  • 利用来自多种模态(包括面部、身体姿势、手势运动和语音)的训练数据,对模型参数进行优化。
  • 最终的情感强度估计值来自各模态在置信度、位移和运动速度上的融合。
  • 采用0–1的唤醒维度量表来表示连续的情感强度水平。

实验结果

研究问题

  • RQ1与单模态方法相比,结合多种模态在情感强度估计方面有何改进?
  • RQ2在面部、身体、手势和语音模态中,哪个模态对情感强度估计准确率的贡献最为显著?
  • RQ3特征点位移和运动速度能否增强实时系统中情感强度的估计?
  • RQ4不同模态的分类置信度水平在加权融合模型中如何相互作用?
  • RQ5所提出的模型在多大程度上与人工标注的情感强度水平保持一致?

主要发现

  • 语音和手势模态在所提出的模型中显著提升了情感强度估计的准确性。
  • 置信度、位移和运动速度的加权和融合方法在性能上优于单一模态。
  • 该模型在0–1唤醒度量表上与人工标注的情感强度水平实现了更好的对齐。
  • 面部和身体模态对模型有所贡献,但其影响程度低于语音和手势模态。
  • 使用Kinect支持的运动追踪技术增强了姿势和手势分析的特征提取。
  • 该模型在多模态人机交互中具备实现实时情感强度估计的可行性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。