[논문 리뷰] Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning
Emotion-LLaMA는 감정 특화 인코더와 지시 튜닝을 통해 오디오, 비주얼, 텍스트 입력을 통합하고 멀티모달 감정 인식 및 추론에서 SOTA를 달성합니다. MERR 데이터셋을 전처训练에 사용하며 EMER, MER2023, DFEW 전반에 걸쳐 제로샷 및 미세조정 성능이 강력함을 보여줍니다.
Accurate emotion perception is crucial for various applications, including human-computer interaction, education, and counseling. However, traditional single-modality approaches often fail to capture the complexity of real-world emotional expressions, which are inherently multimodal. Moreover, existing Multimodal Large Language Models (MLLMs) face challenges in integrating audio and recognizing subtle facial micro-expressions. To address this, we introduce the MERR dataset, containing 28,618 coarse-grained and 4,487 fine-grained annotated samples across diverse emotional categories. This dataset enables models to learn from varied scenarios and generalize to real-world applications. Furthermore, we propose Emotion-LLaMA, a model that seamlessly integrates audio, visual, and textual inputs through emotion-specific encoders. By aligning features into a shared space and employing a modified LLaMA model with instruction tuning, Emotion-LLaMA significantly enhances both emotional recognition and reasoning capabilities. Extensive evaluations show Emotion-LLaMA outperforms other MLLMs, achieving top scores in Clue Overlap (7.83) and Label Overlap (6.25) on EMER, an F1 score of 0.9036 on MER2023-SEMI challenge, and the highest UAR (45.59) and WAR (59.37) in zero-shot evaluations on DFEW dataset.
연구 동기 및 목표
- 단일 모달 접근 방식보다 현실 세계의 다중 모달 맥락에서 정확한 감정 인식을 자극한다.
- 학습을 위한 코스- 및 미세한 감정 주석을 포함하는 크고 다양한 데이터셋(MERR)을 제공한다.
- 감정 특화 인코더와 지시 튜닝을 통해 오디오, 비주얼, 텍스트 입력을 융합하는 Emotion-LLaMA를 제안한다.
- 여러 벤치마크에서 감정 인식 정확도와 추론 능력을 향상시킨다.
제안 방법
- 다양한 감정을 포괄하는 28,618개의 coarse-grained 샘플과 4,487개의 fine-grained 샘플로 MERR 데이터셋 구성.
- 오디오 인코더로 HuBERT를 사용하고, 다중 시각적 특징을 추출하는 다중 뷰 비주얼 인코더(MAE, VideoMAE, EVA)를 활용.
- 학습 가능한 선형 프로젝션을 통해 오디오 및 비주얼 특징을 언어 임베딩 토큰과 공유 공간으로 정렬.
- 지시 튜닝을 적용한 수정된 LLaMA 언어 모델을 도입하여 멀티모달 추론 및 생성을 수행.
- MERR에서의 선행 학습 후 MER2023 및 DFEW에서의 멀티모달 지시 튜닝으로 코스-투-파인 학습을 수행.
실험 결과
연구 질문
- RQ1감정 특화 멀티모달 인코더가 기존 MLLMs보다 인식 및 추론을 향상시킬 수 있는가?
- RQ2다양하고 주석이 달린 멀티모달 감정 데이터셋으로 지시 튜닝이 제로샷 및 미세조정 성능을 향상시키는가?
- RQ3오디오 및 다중 시각 신호가 강건한 감정 이해와 추론에 어떻게 기여하는가?
- RQ4MERR에서의 프리트레이닝이 다른 데이터에 비해 다운스트림 감정 태스크에 미치는 영향은 무엇인가?
주요 결과
- Emotion-LLaMA는 EMER에서 Clue Overlap(7.83) 및 Label Overlap(6.25)에서 최고를 달성합니다.
- MER2023 데이터셋에서 F1 스코어가 0.9036에 도달합니다.
- 제로샷 DFEW 평가에서 최고 UAR(45.59) 및 WAR(59.37)에 도달합니다.
- MER2023에서 A+V+T 모달리티를 사용할 때 F1이 0.9036으로 향상됩니다.
- 비교 전반에 걸쳐 Emotion-LLaMA가 여러 벤치마크에서 다른 MLLMs보다 우수합니다.
- 모델은 총 매개변수의 약 0.495%에 해당하는 34M개의 학습 가능 매개변수만을 사용합니다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.