Skip to main content
QUICK REVIEW

[Paper Review] Self context-aware emotion perception on human-robot interaction

Zihan Lin, Francisco Cruz|arXiv (Cornell University)|Jan 18, 2024
Emotion and Mood RecognitionPsychology3 citations
TL;DR

This paper proposes the Self Context-Aware Model (SCAM), a novel approach for long-term human-robot interaction that improves emotion recognition by integrating contextual emotional continuity using a valence-arousal framework. SCAM enhances accuracy by 9.36% in audio, 3.79% in video, and 1.45% in multimodal settings through a context-aware loss and feature retention mechanism, achieving state-of-the-art performance on the IEMOCAP dataset.

ABSTRACT

Emotion recognition plays a crucial role in various domains of human-robot interaction. In long-term interactions with humans, robots need to respond continuously and accurately, however, the mainstream emotion recognition methods mostly focus on short-term emotion recognition, disregarding the context in which emotions are perceived. Humans consider that contextual information and different contexts can lead to completely different emotional expressions. In this paper, we introduce self context-aware model (SCAM) that employs a two-dimensional emotion coordinate system for anchoring and re-labeling distinct emotions. Simultaneously, it incorporates its distinctive information retention structure and contextual loss. This approach has yielded significant improvements across audio, video, and multimodal. In the auditory modality, there has been a notable enhancement in accuracy, rising from 63.10% to 72.46%. Similarly, the visual modality has demonstrated improved accuracy, increasing from 77.03% to 80.82%. In the multimodal, accuracy has experienced an elevation from 77.48% to 78.93%. In the future, we will validate the reliability and usability of SCAM on robots through psychology experiments.

Motivation & Objective

  • To address the limitation of existing emotion recognition models that ignore emotional continuity in long-term human-robot interactions.
  • To improve emotion recognition accuracy by incorporating the robot’s own emotional context and prior emotional states.
  • To bridge discrete and continuous emotion models by anchoring emotions in a valence-arousal space for better temporal consistency.
  • To validate the effectiveness of contextual loss and feature retention in enhancing multimodal emotion perception.
  • To lay the foundation for future psychological validation of SCAM in real robotic systems.

Proposed method

  • SCAM employs a two-dimensional valence-arousal coordinate system to anchor and re-label emotions, enabling the model to learn relationships between basic and non-basic emotions.
  • It introduces a novel contextual loss that encourages the model to predict the emotion of preceding segments, capturing emotional trend continuity.
  • The model uses a dedicated information retention structure to preserve and integrate relevant features from prior emotional contexts during prediction.
  • SCAM combines unimodal and multimodal learning with a multitask loss that jointly optimizes for emotion classification, valence, and arousal prediction.
  • The framework applies relabeling based on valence and arousal to refine emotion representations and improve generalization across modalities.
  • It leverages the IEMOCAP dataset for training and evaluation across audio, visual, and multimodal configurations.
Figure 1: Context interaction in HRI
Figure 1: Context interaction in HRI

Experimental results

Research questions

  • RQ1Can integrating emotional context from prior interactions significantly improve long-term emotion recognition in human-robot interaction?
  • RQ2How does modeling emotional continuity through valence and arousal enhance recognition accuracy across different modalities?
  • RQ3To what extent does the proposed contextual loss improve the model’s ability to predict emotional trends over time?
  • RQ4How does SCAM compare to baseline models in handling conflicting or ambiguous multimodal signals?
  • RQ5Can the integration of continuous and discrete emotion modeling lead to more robust and accurate emotion perception?

Key findings

  • SCAM achieved a 9.36% improvement in emotion recognition accuracy on the auditory modality, rising from 63.10% to 72.46%.
  • The visual modality saw a 3.79% accuracy gain, increasing from 77.03% to 80.82%.
  • In the multimodal setting, accuracy improved by 1.45%, from 77.48% to 78.93%.
  • The context loss component consistently decreased during training, indicating effective learning of emotional continuity despite fluctuations in total test loss.
  • Visualization confirmed that SCAM maintains accurate emotion prediction even when emotional context changes continuously, independent of label consistency.
  • Error analysis revealed that auditory and visual modalities make complementary errors, with visual modality outperforming audio in identifying 'happy' and 'neutral' emotions.
Figure 2: IEMOCAP emotions on Valence-Arousal axis
Figure 2: IEMOCAP emotions on Valence-Arousal axis

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.