Skip to main content
QUICK REVIEW

[논문 리뷰] Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning

Zebang Cheng, Zhi-Qi Cheng|arXiv (Cornell University)|2024. 06. 17.
Emotion and Mood Recognition인용 수 10
한 줄 요약

Emotion-LLaMA는 감정 특화 인코더와 지시 튜닝을 통해 오디오, 비주얼, 텍스트 입력을 통합하고 멀티모달 감정 인식 및 추론에서 SOTA를 달성합니다. MERR 데이터셋을 전처训练에 사용하며 EMER, MER2023, DFEW 전반에 걸쳐 제로샷 및 미세조정 성능이 강력함을 보여줍니다.

ABSTRACT

Accurate emotion perception is crucial for various applications, including human-computer interaction, education, and counseling. However, traditional single-modality approaches often fail to capture the complexity of real-world emotional expressions, which are inherently multimodal. Moreover, existing Multimodal Large Language Models (MLLMs) face challenges in integrating audio and recognizing subtle facial micro-expressions. To address this, we introduce the MERR dataset, containing 28,618 coarse-grained and 4,487 fine-grained annotated samples across diverse emotional categories. This dataset enables models to learn from varied scenarios and generalize to real-world applications. Furthermore, we propose Emotion-LLaMA, a model that seamlessly integrates audio, visual, and textual inputs through emotion-specific encoders. By aligning features into a shared space and employing a modified LLaMA model with instruction tuning, Emotion-LLaMA significantly enhances both emotional recognition and reasoning capabilities. Extensive evaluations show Emotion-LLaMA outperforms other MLLMs, achieving top scores in Clue Overlap (7.83) and Label Overlap (6.25) on EMER, an F1 score of 0.9036 on MER2023-SEMI challenge, and the highest UAR (45.59) and WAR (59.37) in zero-shot evaluations on DFEW dataset.

연구 동기 및 목표

  • 단일 모달 접근 방식보다 현실 세계의 다중 모달 맥락에서 정확한 감정 인식을 자극한다.
  • 학습을 위한 코스- 및 미세한 감정 주석을 포함하는 크고 다양한 데이터셋(MERR)을 제공한다.
  • 감정 특화 인코더와 지시 튜닝을 통해 오디오, 비주얼, 텍스트 입력을 융합하는 Emotion-LLaMA를 제안한다.
  • 여러 벤치마크에서 감정 인식 정확도와 추론 능력을 향상시킨다.

제안 방법

  • 다양한 감정을 포괄하는 28,618개의 coarse-grained 샘플과 4,487개의 fine-grained 샘플로 MERR 데이터셋 구성.
  • 오디오 인코더로 HuBERT를 사용하고, 다중 시각적 특징을 추출하는 다중 뷰 비주얼 인코더(MAE, VideoMAE, EVA)를 활용.
  • 학습 가능한 선형 프로젝션을 통해 오디오 및 비주얼 특징을 언어 임베딩 토큰과 공유 공간으로 정렬.
  • 지시 튜닝을 적용한 수정된 LLaMA 언어 모델을 도입하여 멀티모달 추론 및 생성을 수행.
  • MERR에서의 선행 학습 후 MER2023 및 DFEW에서의 멀티모달 지시 튜닝으로 코스-투-파인 학습을 수행.

실험 결과

연구 질문

  • RQ1감정 특화 멀티모달 인코더가 기존 MLLMs보다 인식 및 추론을 향상시킬 수 있는가?
  • RQ2다양하고 주석이 달린 멀티모달 감정 데이터셋으로 지시 튜닝이 제로샷 및 미세조정 성능을 향상시키는가?
  • RQ3오디오 및 다중 시각 신호가 강건한 감정 이해와 추론에 어떻게 기여하는가?
  • RQ4MERR에서의 프리트레이닝이 다른 데이터에 비해 다운스트림 감정 태스크에 미치는 영향은 무엇인가?

주요 결과

  • Emotion-LLaMA는 EMER에서 Clue Overlap(7.83) 및 Label Overlap(6.25)에서 최고를 달성합니다.
  • MER2023 데이터셋에서 F1 스코어가 0.9036에 도달합니다.
  • 제로샷 DFEW 평가에서 최고 UAR(45.59) 및 WAR(59.37)에 도달합니다.
  • MER2023에서 A+V+T 모달리티를 사용할 때 F1이 0.9036으로 향상됩니다.
  • 비교 전반에 걸쳐 Emotion-LLaMA가 여러 벤치마크에서 다른 MLLMs보다 우수합니다.
  • 모델은 총 매개변수의 약 0.495%에 해당하는 34M개의 학습 가능 매개변수만을 사용합니다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.