Skip to main content
QUICK REVIEW

[논문 리뷰] Multiple Appropriate Facial Reaction Generation in Dyadic Interaction Settings: What, Why and How?

Siyang Song, Micol Spitale|arXiv (Cornell University)|2023. 02. 13.
Color perception and design인용 수 5
한 줄 요약

이 논문은 이인 상호작용에서 다수의 적절한 얼굴 반응 생성(Multiple Appropriate Facial Reaction Generation, fMARG)이라는 새로운 작업을 제안하며, 딥러닝을 활용해 맥락에 적절한 다수의 얼굴 반응을 예측하고 생성하며 평가하는 프레임워크를 제시한다. 정확도, 다양성, 현실성, 동기화성이라는 첫 번째 객관적 평가 지표를 수립하여, 동일한 화자 행동에 대해 다양한 맥락에서 다양한, 현실적인, 시간적으로 정렬된 얼굴 반응을 생성할 수 있음을 입증한다.

ABSTRACT

According to the Stimulus Organism Response (SOR) theory, all human behavioral reactions are stimulated by context, where people will process the received stimulus and produce an appropriate reaction. This implies that in a specific context for a given input stimulus, a person can react differently according to their internal state and other contextual factors. Analogously, in dyadic interactions, humans communicate using verbal and nonverbal cues, where a broad spectrum of listeners' non-verbal reactions might be appropriate for responding to a specific speaker behaviour. There already exists a body of work that investigated the problem of automatically generating an appropriate reaction for a given input. However, none attempted to automatically generate multiple appropriate reactions in the context of dyadic interactions and evaluate the appropriateness of those reactions using objective measures. This paper starts by defining the facial Multiple Appropriate Reaction Generation (fMARG) task for the first time in the literature and proposes a new set of objective evaluation metrics to evaluate the appropriateness of the generated reactions. The paper subsequently introduces a framework to predict, generate, and evaluate multiple appropriate facial reactions.

연구 동기 및 목표

  • 이재료 문헌에서 다수의 적절한 비언어적 반응 생성을 다루는 다수의 적절한 얼굴 반응 생성(fMARG) 작업을 처음으로 정의함.
  • 한 명의 화자 행동에 대해 맥락과 개인차를 고려하여 다수의 적절한 얼굴 반응을 예측하고 생성하며 평가하는 프레임워크를 개발함.
  • 정확도, 다양성, 현실성, 동기화성이라는 객관적 평가 지표 세트를 도입하여 주관적 인간 평가에 의존하지 않고도 생성된 얼굴 반응의 품질을 평가함.
  • 유사도 임계값과 실제 반응 분포 기반의 맥락 인식 전략을 사용하여 적절한 얼굴 반응의 자동 레이블링을 가능하게 함.
  • 프레임워크가 이인 상호작용의 인간 행동 규범을 반영하며, 다양한, 현실적인, 시간적으로 동기화된 얼굴 반응을 생성할 수 있음을 검증함.

제안 방법

  • 프레임워크는 딥러닝 기반 모델을 사용하여 단일 화자 행동에서 맥락과 청취자 특성 상태에 조건화된 다수의 얼굴 반응(α개)을 생성함.
  • 유사도 임계값(T)을 사용하여 생성된 반응이 적절한지 판단함. 동일 맥락 내에서 생성된 반응과 실제 얼굴 반응 간의 코사인 유사도를 기반으로 함.
  • 정확도 지표는 데이터셋 내 적절한 실제 반응 중 하나 이상과 유사도가 높은(임계값 초과) 생성 반응의 비율로 계산됨.
  • 다양성은 얼굴 반응 분산(FRVar), 생성 반응 간 MSE 합계(S-MSE), 조건 간 다양성(FRDvs)을 통해 측정되어 다양한 출력을 보장함.
  • 현실성은 생성된 반응 분포와 실제 반응 분포 간의 프레셰 인ception 거리(FID)를 사용하여 평가됨.
  • 동기화성은 화자 행동과 생성된 얼굴 반응 시계열 간의 시간 지연 교차상관관계(TLCC)를 사용하여 평가되어 시간적 정렬을 확보함.
Figure 1: Multiple appropriate reaction generation in dyadic interaction settings. The same input stimuli ( $(b(S_{1})^{t})_{1}=(b(S_{1})^{t})_{m}$ ) under different contexts ( $C_{1}\neq C_{m}$ ) of the speaker $S_{1}$ can elicit different reactions for the same listener ( $(f_{1})_{1}\neq(f_{1})_{
Figure 1: Multiple appropriate reaction generation in dyadic interaction settings. The same input stimuli ( $(b(S_{1})^{t})_{1}=(b(S_{1})^{t})_{m}$ ) under different contexts ( $C_{1}\neq C_{m}$ ) of the speaker $S_{1}$ can elicit different reactions for the same listener ( $(f_{1})_{1}\neq(f_{1})_{

실험 결과

연구 질문

  • RQ1이인 상호작용에서 적절한 얼굴 반응은 무엇으로 정의되며, 동일한 화자 행동에 대해 다수의 적절한 반응이 동시에 존재할 수 있는가?
  • RQ2어떻게 맥락 인식 기반으로 생성된 얼굴 반응의 적절성, 다양성, 현실성, 동기화성을 객관적으로 평가할 수 있는가?
  • RQ3딥러닝 프레임워크는 단일 화자 행동에 대해 다수의 구분 가능한, 현실적인, 시간적으로 동기화된 얼굴 반응을 생성할 수 있는가?
  • RQ4맥락적 요소와 개인적 차이는 이인 상호작용에서 적절한 얼굴 반응 생성에 어떻게 영향을 미치는가?
  • RQ5주관적 인간 평가에 의존하지 않고도 다수의 얼굴 반응 생성 품질을 신뢰성 있게 평가할 수 있는 객관적 지표는 무엇인가?

주요 결과

  • 제안된 fMARG 프레임워크는 입력 화자 행동당 다수의 구분 가능한 얼굴 반응을 성공적으로 생성하여 생성 출력 간 높은 다양성을 보임.
  • 정확도 지표는 생성된 반응의 상당 부분이 실제 적절한 반응과 유사함을 보이며, 임계값 기반의 유사도 전략이 적절성 평가의 신뢰성을 확보함.
  • 다양성 지표인 FRVar, S-MSE, FRDvs는 입력 조건 내외에서 다양한 얼굴 반응을 생성함을 확인함.
  • FID로 측정된 현실성 점수는 생성된 얼굴 반응이 분포적으로 실제 인간 반응과 시각적·통계적으로 유사함을 나타냄.
  • 동기화성 지표(FRSyn)는 생성된 얼굴 반응이 화자 행동과 시간적으로 정렬되어 있음을 확인하여 자연스러운 상호작용 역학을 반영함.
  • 프레임워크는 객관적이고 재현 가능한 지표를 사용하여 얼굴 반응 생성 평가의 새로운 기준을 설정함으로써 향후 정서 컴퓨팅 및 인간-로봇 상호작용 분야의 연구를 가능하게 함.
Figure 2: Illustration of the proposed automatic appropriate facial reaction labelling strategy ( Sec. 3.1 ). It first represents each audio-facial speaker behaviour clip as a multi-channel time-series signal. Then, we compute the similarity between each pair of speaker behaviour time-series ( Step
Figure 2: Illustration of the proposed automatic appropriate facial reaction labelling strategy ( Sec. 3.1 ). It first represents each audio-facial speaker behaviour clip as a multi-channel time-series signal. Then, we compute the similarity between each pair of speaker behaviour time-series ( Step

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.