Skip to main content
QUICK REVIEW

[논문 리뷰] Audio-Visual Fusion for Emotion Recognition in the Valence-Arousal Space Using Joint Cross-Attention

R. Gnana Praveen, Éric Granger|arXiv (Cornell University)|2022. 09. 19.
Emotion and Mood Recognition인용 수 4
한 줄 요약

이 논문은 차원적 정서 인식에서 청각-시각 융합을 위한 공동 교차주의(Joint Cross-Attention, JCA) 메커니즘을 제안한다. 이 메커니즘은 공동 및 개별 모odal 특징 간 상관관계를 기반으로 한 주의 힘을 계산하여 내모달 및 간모달 관계를 모두 모델링한다. 제안된 방법은 RECOLA 및 AffWild2에서 최신 기술보다 뛰어난 성능을 보이며, 특히 소음이 많거나 청각 데이터가 누락된 조건에서도 강건성을 입증하고, 정서의 진동성과 각성도 예측에서 향상된 성능을 보인다.

ABSTRACT

Automatic emotion recognition (ER) has recently gained lot of interest due to its potential in many real-world applications. In this context, multimodal approaches have been shown to improve performance (over unimodal approaches) by combining diverse and complementary sources of information, providing some robustness to noisy and missing modalities. In this paper, we focus on dimensional ER based on the fusion of facial and vocal modalities extracted from videos, where complementary audio-visual (A-V) relationships are explored to predict an individual's emotional states in valence-arousal space. Most state-of-the-art fusion techniques rely on recurrent networks or conventional attention mechanisms that do not effectively leverage the complementary nature of A-V modalities. To address this problem, we introduce a joint cross-attentional model for A-V fusion that extracts the salient features across A-V modalities, that allows to effectively leverage the inter-modal relationships, while retaining the intra-modal relationships. In particular, it computes the cross-attention weights based on correlation between the joint feature representation and that of the individual modalities. By deploying the joint A-V feature representation into the cross-attention module, it helps to simultaneously leverage both the intra and inter modal relationships, thereby significantly improving the performance of the system over the vanilla cross-attention module. The effectiveness of our proposed approach is validated experimentally on challenging videos from the RECOLA and AffWild2 datasets. Results indicate that our joint cross-attentional A-V fusion model provides a cost-effective solution that can outperform state-of-the-art approaches, even when the modalities are noisy or absent.

연구 동기 및 목표

  • 비디오 데이터로부터 청각 및 시각 모달을 효과적으로 융합하여 정서의 진동성-각성도 공간에서의 차원적 정서 인식을 향상시키는 것.
  • 기존 융합 방법이 청각-시각 특징의 내모달 및 간모달 관계를 공동으로 모델링하지 못하는 한계를 해결하는 것.
  • 청각 모달이 소음이 있거나 누락된 경우에도 높은 성능을 유지할 수 있는 비용 효율적이고 강건한 융합 메커니즘을 개발하는 것.
  • 다양한 정서 표현과 도전적인 조건을 가진 실제 데이터셋에서 제안된 방법을 검증하는 것.

제안 방법

  • 청각(A) 및 시각(V) 모달에 대해 별도의 백본을 사용하여 초기 특징을 추출한다.
  • 공동 A-V 특징 표현과 개별 A 및 V 특징 간 상관관계를 기반으로 공동 교차주의 주의 힘을 계산한다.
  • 주의 메커니즘은 각 모달이 다른 모달뿐만 아니라 자신의 내모달 표현에도 주의를 기울일 수 있도록 하여 간모달 및 내모달 관계를 모두 유지한다.
  • 양 모달의 주의를 받은 특징을 연결하여 완전히 연결된 레이어를 통과시켜 진동성 및 각성도 점수를 예측한다.
  • 모델은 엔드 투 엔드로 훈련되며, 청각 세그먼트가 누락된 경우를 포함한 다양한 조건에서 평가된다.
  • 이 방법은 모듈식이며, A 및 V 모달에 대해 다양한 백본 아키텍처와 호환되도록 설계되어 있다.
Figure 1: The valence-arousal space. Valence denotes the range of emotions from being very sad (negative) to very happy (positive) and arousal reflects the energy or intensity of emotions from very passive to very active.
Figure 1: The valence-arousal space. Valence denotes the range of emotions from being very sad (negative) to very happy (positive) and arousal reflects the energy or intensity of emotions from very passive to very active.

실험 결과

연구 질문

  • RQ1내모달 및 간모달 관계를 공동으로 모델링하면 청각-시각 융합의 정성적 정서 인식 성능이 향상되는가?
  • RQ2공동 및 개별 모달 특징 간 상관관계 기반 주의는 일반 교차주의에 비해 특징 표현을 어떻게 향상시키는가?
  • RQ3청각 모달이 소음이 있거나 존재하지 않을 경우 제안된 방법이 얼마나 높은 성능을 유지하는가?
  • RQ4정서 예측 과정에서 시간적 변화, 가림, 표정 자세 변화 등의 영향을 모델은 어떻게 처리하는가?
  • RQ5공동 교차주의 메커니즘은 이전의 주의 기반 융합 방법에 비해 더 정확하고 맥락을 고려한 주의 맵을 도출하는가?

주요 결과

  • 제안된 공동 교차주의(Joint Cross-Attention, JCA) 모델은 RECOLA 및 AffWild2 데이터셋에서 진동성 및 각성도 예측에서 최신 기술보다 뛰어난 성능을 보였다.
  • 청각 세그먼트가 누락되거나 소음이 있는 경우에도 뛰어난 성능을 유지하여 모달 품질 열화에 대한 강건성을 입증했다.
  • 시각화 결과 JCA는 일반 교차주의가 자주 핵심 시간 클립을 놓치는 것과 달리, 관련 있는 얼굴 표정과 음성 에너지 변화에 효과적으로 집중하는 것으로 나타났다.
  • 특히 표정 자세 변화나 부분적인 얼굴 가림 상황에서도 실제 진동성 및 각성도를 더 잘 추적하는 경향을 보였다.
  • 강력한 성능에도 불구하고, 뿌연 이미지나 자세 변화로 인해 주의 메커니즘이 잘못 이끌릴 경우 각성도 예측의 정확도는 진동성 예측보다 약간 낮게 나타났다.
  • 주의 메커니즘이 청각 및 시각 모달 간의 맥락적 관계를 성공적으로 포착하여 더 정확하고 시간적으로 일관된 예측을 이끌어냈다.
Figure 2: Joint cross-attention model proposed for A-V fusion (in testing mode).
Figure 2: Joint cross-attention model proposed for A-V fusion (in testing mode).

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.