Skip to main content
QUICK REVIEW

[논문 리뷰] IDRL: An Individual-Aware Multimodal Depression-Related Representation Learning Framework for Depression Diagnosis

Chongxiao Wang, Junjie Liang|arXiv (Cornell University)|2026. 03. 12.
Emotion and Mood Recognition인용 수 0
한 줄 요약

IDRL은 멀티모달 우울 신호를 공통 공간, 특정 공간, 무관한 공간으로 분리하고 개별 인식 융합을 통해 모달리티 간의 견고한 우울 진단을 위한 특성 가중치를 적응적으로 조정합니다.

ABSTRACT

Depression is a severe mental disorder, and reliable identification plays a critical role in early intervention and treatment. Multimodal depression detection aims to improve diagnostic performance by jointly modeling complementary information from multiple modalities. Recently, numerous multimodal learning approaches have been proposed for depression analysis; however, these methods suffer from the following limitations: 1) inter-modal inconsistency and depression-unrelated interference, where depression-related cues may conflict across modalities while substantial irrelevant content obscures critical depressive signals, and 2) diverse individual depressive presentations, leading to individual differences in modality and cue importance that hinder reliable fusion. To address these issues, we propose Individual-aware Multimodal Depression-related Representation Learning Framework (IDRL) for robust depression diagnosis. Specifically, IDRL 1) disentangles multimodal representations into a modality-common depression space, a modality-specific depression space, and a depression-unrelated space to enhance modality alignment while suppressing irrelevant information, and 2) introduces an individual-aware modality-fusion module (IAF) that dynamically adjusts the weights of disentangled depression-related features based on their predictive significance, thereby achieving adaptive cross-modal fusion for different individuals. Extensive experiments demonstrate that IDRL achieves superior and robust performance for multimodal depression detection.

연구 동기 및 목표

  • 다중 모달 간 차이와 개인 표현 차이에도 불구하고 로버스트한 멀티모달 우울 탐지를 촉진한다.
  • 모달리티 공통 정보, 모달리티 특이 정보, 그리고 우울과 무관한 정보를 분리하는 프레임워크를 제안한다.
  • 개별 인식 융합 메커니즘을 도입하여 각 개인별로 특징에 가중치를 적응적으로 부여한다.
  • ABL-에서 벤치마크 데이터셋 AVEC-2014 및 Twitter을 대상으로 제거/가시화와 함께 효과를 검증한다.

제안 방법

  • 모달리티별 인코더를 이용하여 멀티모달 표현을 모달리티-공통(F_c^m), 모달리티-특이(F_s^m), 및 우울-무관(N_c^m, N_s^m) 공간으로 분리한다.
  • 정보 보존 및 교차 모달 상호작용을 강제하기 위해 자기- 및 교차 모달 재구성을 통해 원래 특징을 재구성한다.
  • Central Moment Discrepancy(CMD)를 적용하여 모달리티-공통 특징을 모달리티 간에 정렬한다.
  • 분리된 공간 간의 분리를 유도하기 위한 소프트 직교 정규화를 활용한다.
  • concat된 특징에 대해 개인 인식 주의 기반 융합을 통해 F_S를 예측하는 융합 표현을 사용한다.
  • 정보가 풍부한 융합과 예측 중요도와의 일관성을 촉진하기 위해 보조 기여도 및 정렬 손실을 도입한다.
  • 진단, 분리, 개인 인식 구성 요소를 지정된 가중치로 결합한 총 손실을 최적화한다.

실험 결과

연구 질문

  • RQ1모달리티 공통 정보, 모달리티 특이 정보, 그리고 우울-무관 정보를 구분해도 교차 모달 우울 탐지 성능이 향상되는가?
  • RQ2개인 인식 융합 모듈이 서로 다른 우울 표현을 가진 개인들 간의 적응적 멀티모달 융합을 개선하는가?
  • RQ3제안된 손실 구성 요소가 모델 성능과 특징 분리에 어떤 영향을 미치는가?
  • RQ4제안된 방법이 비디오/오디오 및 텍스트/이미지의 다양한 모달리티 페어와 데이터셋에서 일반화되는가?

주요 결과

  • IDRL은 AVEC-2014의 비디오+오디오에서, Twitter의 텍스트+이미지에서 최첨단 성능을 달성한다.
  • 모달리티를 공통, 특이, 무관 공간으로 분리하면 간섭이 감소하고 정렬이 개선된다.
  • 개인 인식 융합은 비적응적 융합보다 더 나은 성능으로 이어지는 적응적 가중치를 제공한다.
  • 아블레이션 연구에서 직교성 및 CMD 손실이 성능과 분리 품질에 결정적임이 확인된다.
  • 시각화(t-SNE, Grad-CAM++)은 전체 모델을 사용할 때 특징 공간의 분리가 더 명확하고 예측적 단서가 더 집중됨을 보여준다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.