Skip to main content
QUICK REVIEW

[论文解读] Self context-aware emotion perception on human-robot interaction

Zihan Lin, Francisco Cruz|arXiv (Cornell University)|Jan 18, 2024
Emotion and Mood RecognitionPsychology被引用 3
一句话总结

本文提出了一种自洽上下文感知模型(SCAM),这是一种新颖的长期人机交互方法,通过采用效价-唤醒框架整合上下文情感连续性,提升了情感识别性能。SCAM通过上下文感知损失和特征保留机制,在音频、视频和多模态设置下分别实现了9.36%、3.79%和1.45%的准确率提升,在IEMOCAP数据集上达到了最先进性能。

ABSTRACT

Emotion recognition plays a crucial role in various domains of human-robot interaction. In long-term interactions with humans, robots need to respond continuously and accurately, however, the mainstream emotion recognition methods mostly focus on short-term emotion recognition, disregarding the context in which emotions are perceived. Humans consider that contextual information and different contexts can lead to completely different emotional expressions. In this paper, we introduce self context-aware model (SCAM) that employs a two-dimensional emotion coordinate system for anchoring and re-labeling distinct emotions. Simultaneously, it incorporates its distinctive information retention structure and contextual loss. This approach has yielded significant improvements across audio, video, and multimodal. In the auditory modality, there has been a notable enhancement in accuracy, rising from 63.10% to 72.46%. Similarly, the visual modality has demonstrated improved accuracy, increasing from 77.03% to 80.82%. In the multimodal, accuracy has experienced an elevation from 77.48% to 78.93%. In the future, we will validate the reliability and usability of SCAM on robots through psychology experiments.

研究动机与目标

  • 解决现有情感识别模型在长期人机交互中忽略情感连续性的问题。
  • 通过整合机器人自身的情感上下文和先前情感状态,提升情感识别准确率。
  • 通过将情感锚定在效价-唤醒空间,弥合离散与连续情感模型之间的鸿沟,实现更好的时间一致性。
  • 验证上下文损失和特征保留机制在提升多模态情感感知方面的有效性。
  • 为未来在真实机器人系统中对SCAM进行心理层面验证奠定基础。

提出的方法

  • SCAM采用二维效价-唤醒坐标系对情感进行锚定和重标记,使模型能够学习基本情感与非基本情感之间的关系。
  • 提出一种新颖的上下文损失,鼓励模型预测前序片段的情感,从而捕捉情感趋势的连续性。
  • 模型采用专用的信息保留结构,在预测过程中保留并整合先前情感上下文的相关特征。
  • SCAM结合单模态与多模态学习,通过多任务损失联合优化情感分类、效价和唤醒预测。
  • 框架基于效价和唤醒对情感表示进行重标记,以提升跨模态的泛化能力。
  • 采用IEMOCAP数据集进行训练与评估,覆盖音频、视觉和多模态配置。
Figure 1: Context interaction in HRI
Figure 1: Context interaction in HRI

实验结果

研究问题

  • RQ1将先前交互中的情感上下文整合进来,是否能显著提升长期人机交互中的情感识别性能?
  • RQ2通过效价和唤醒建模情感连续性,如何提升不同模态下的识别准确率?
  • RQ3所提出的上下文损失在多大程度上提升了模型预测随时间变化的情感趋势的能力?
  • RQ4SCAM在处理冲突或模糊的多模态信号时,与基线模型相比表现如何?
  • RQ5将连续与离散情感建模相结合,是否能带来更鲁棒且更准确的情感感知?

主要发现

  • 在音频模态上,SCAM的情感识别准确率提升了9.36%,从63.10%提高到72.46%。
  • 在视觉模态上,准确率提升了3.79%,从77.03%提高到80.82%。
  • 在多模态设置中,准确率提升了1.45%,从77.48%提高到78.93%。
  • 上下文损失组件在训练过程中持续下降,表明尽管总测试损失存在波动,模型仍有效学习了情感连续性。
  • 可视化结果表明,即使情感上下文持续变化,SCAM仍能保持准确的情感预测,且不受标签一致性的影响。
  • 错误分析显示,音频与视觉模态的错误具有互补性,视觉模态在识别‘快乐’和‘中性’情感方面优于音频模态。
Figure 2: IEMOCAP emotions on Valence-Arousal axis
Figure 2: IEMOCAP emotions on Valence-Arousal axis

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。