[논문 리뷰] HCAM -- Hierarchical Cross Attention Model for Multi-modal Emotion Recognition
HCAM은 co-attention 모듈을 통해 오디오(wav2vec)와 텍스트(RoBERTa)를 융합하는 계층적 다중 모달 프레임워크를 제시하며, 감독된 대조 학습(loss)을 세 단계로 훈련하여 IEMOCAP, MELD, CMU-MOSI에서 최첨단 성과를 달성한다.
Emotion recognition in conversations is challenging due to the multi-modal nature of the emotion expression. We propose a hierarchical cross-attention model (HCAM) approach to multi-modal emotion recognition using a combination of recurrent and co-attention neural network models. The input to the model consists of two modalities, i) audio data, processed through a learnable wav2vec approach and, ii) text data represented using a bidirectional encoder representations from transformers (BERT) model. The audio and text representations are processed using a set of bi-directional recurrent neural network layers with self-attention that converts each utterance in a given conversation to a fixed dimensional embedding. In order to incorporate contextual knowledge and the information across the two modalities, the audio and text embeddings are combined using a co-attention layer that attempts to weigh the utterance level embeddings relevant to the task of emotion recognition. The neural network parameters in the audio layers, text layers as well as the multi-modal co-attention layers, are hierarchically trained for the emotion classification task. We perform experiments on three established datasets namely, IEMOCAP, MELD and CMU-MOSI, where we illustrate that the proposed model improves significantly over other benchmarks and helps achieve state-of-art results on all these datasets.
연구 동기 및 목표
- 대화에서 다중 모달 신호(오디오 및 텍스트)를 활용하여 감정 인식을 다룬다.
- 먼저 단일 모달 발화 임베딩을 학습하고, 그다음 컨텍스트를 통합하며, 마지막으로 교차 모달 융합을 수행하는 계층적 학습 파이프라인을 도입한다.
- 교차 엔트로피 외에 감독 대조 손실을 도입하여 어려운 예제 학습을 향상시키는 것을 탐구한다.
- IEMOCAP, MELD, CMU-MOSI에서 접근법을 평가하고 베이스라인 및 기존의 최첨단보다 개선을 입증한다.
- 텍스트 모달리티에 대한 ASR 전사에 대한 강건성을 평가한다.
제안 방법
- 학습 가능한 wav2vec 2.0 기반 프런트 엔드를 사용하고 프레임 전체에 걸쳐 트랜스포머 계층 출력을 합산한 다음, 1-D CNN과 발화 수준 풀링을 적용해 오디오 임베딩을 생성한다.
- RoBERTa 임베딩(마지막 4개 층)과 Bi-GRU를 사용해 발화 수준 텍스트 임베딩을 형성하고 전체 파인튜닝을 수행한다.
- Stage II는 각 모달리티의 발화 간 맥락(context)을 인코딩하기 위해 self-attention을 갖춘 contextual-GRU를 사용한다.
- Stage III는 cross-attention 및 self-attention 블록으로 구성된 공동 주의 네트워크를 통해 모달리티를 융합하고, 결합된 다중 모달 표현을 생성한다.
- Ltot = beta·CE + (1−beta)·Lsup-con로 합성 손실로 학습하며, Lsup-con은 배치 내 발화들 간의 감독 대조 손실이다.
- 추론 시 오디오, 텍스트, 그리고 co-attention 모듈의 스테이지별 출력을 가중 평균해 예측을 앙상블한다.

실험 결과
연구 질문
- RQ1 hierarchical, stage-wise training regime가 ERC를 개별 발화 수준, 맥락, 교차 모달 정보를 점진적으로 도입하여 개선할 수 있는가?
- RQ2co-attention이 다양한 데이터셋에서 대화의 감정 인식에 대해 오디오와 텍스트 모듈을 효과적으로 융합하는가?
- RQ3감독 대조 손실 도입이 ERC 작업의 학습 및 성능을 개선하는가?
- RQ4테스트 시 ASR 기반 전사에 HCAM의 강건성은 어느 정도인가?
주요 결과
- Stage I (audio) 은 IEMOCAP에서 64.3% (4-way) 및 64.4% (table의 Stage I 텍스트 정렬 포함, 4-way) 가중 F1을 달성한다; TODAY 값은 각 모듈에 대해 논문이 보고한 수치를 반영한다.
- Stage II (맥락적 오디오) 는 IEMOCAP 4-way에서 78.7%; IEMOCAP 6-way에서 65.7%; 데이터 세트 전반에 걸쳐 Stage I 대비 오디오의 성능이 +14~+19 포인트 개선된다.
- Stage I (텍스트) 은 RoBERTa 기반 임베딩으로 미세조정 후 IEMOCAP 4-way에서 68.4%; CMU-MOSI에서 84.3%를 달성한다. Stage II 텍스트는 81.4% (IEMOCAP 4-way); 85.4% (CMU-MOSI)이다.
- Stage III 다중 모달 융합 (오디오+텍스트)은 IEMOCAP 4-way에서 85.9%; IEMOCAP 6-way에서 70.5%; MELD 7-way에서 65.8%; CMU-MOSI에서 85.8%를 달성한다.
- HCAM은 이전 연구들에 비해 IEMOCAP, MELD, CMU-MOSI에서 최첨단 다중 모달 융합 성능을 달성한다.
- 테스트 시 ASR 전사는 원본 텍스트 결과에 비해 IEMOCAP 4-way에서 84.6%, CMU-MOSI에서 65.8%를 산출하며, 일부 설정에서 약간의 하락이 있다.

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.