[논문 리뷰] MIMAMO Net: Integrating Micro- and Macro-motion for Video Emotion Recognition
이 논문은 상반된 위상 차이를 광학 흐름 대신 시간적 입력으로 사용하여 미세운동 및 거시운동 특징을 통합함으로써 비디오 정서 인식을 위한 이중 스트림 순환 신경망인 MIMAMO Net을 제안한다. 이 모델은 OMG-Emotion 및 Aff-Wild 데이터셋에서 최신 기술 수준의 성능을 달성하며, 자극성 예측에서 뚜렷한 향상과 함께 조명 변화에 대한 강건성도 향상시킨다.
Spatial-temporal feature learning is of vital importance for video emotion recognition. Previous deep network structures often focused on macro-motion which extends over long time scales, e.g., on the order of seconds. We believe integrating structures capturing information about both micro- and macro-motion will benefit emotion prediction, because human perceive both micro- and macro-expressions. In this paper, we propose to combine micro- and macro-motion features to improve video emotion recognition with a two-stream recurrent network, named MIMAMO (Micro-Macro-Motion) Net. Specifically, smaller and shorter micro-motions are analyzed by a two-stream network, while larger and more sustained macro-motions can be well captured by a subsequent recurrent network. Assigning specific interpretations to the roles of different parts of the network enables us to make choice of parameters based on prior knowledge: choices that turn out to be optimal. One of the important innovations in our model is the use of interframe phase differences rather than optical flow as input to the temporal stream. Compared with the optical flow, phase differences require less computation and are more robust to illumination changes. Our proposed network achieves state of the art performance on two video emotion datasets, the OMG emotion dataset and the Aff-Wild dataset. The most significant gains are for arousal prediction, for which motion information is intuitively more informative. Source code is available at https://github.com/wtomin/MIMAMO-Net.
연구 동기 및 목표
- 얼굴 표정의 미세운동 및 거시운동을 동시에 모델링하여 비디오 정서 인식 성능을 향상시키기.
- 조명 변화에 취약하고 계산 비용이 높은 광학 흐름의 한계를 해결하기.
- 복합 기울일 수 있는 피라미드에서 유도된 위상 차이를 활용해 더 강건하고 효율적인 운동 표현 방식을 개발하기.
- 일반적으로 간과되기 쉬운 미세운동 특징이 자극성 예측에 기여하는 바가 크다는 것을 입증하기.
- 시간 스트림 입력으로서 위상 차이가 광학 흐름이나 원시 위상 이미지보다 우수한가를 검증하기.
제안 방법
- MIMAMO Net은 외관 특징을 위한 공간 스트림과 운동 특징을 위한 시간 스트림을 갖는 이중 스트림 아키텍처를 사용한다.
- 시간 스트림은 광학 흐름을 대체하여 복소 기울일 수 있는 피라미드에서 유도된 프레임 간 위상 차이를 운동 표현으로 사용한다.
- 위상 차이는 기울일 수 푸리에 변환의 연속된 프레임 간 위상 이동으로 계산되며, 조명 불변성과 낮은 계산 비용을 제공한다.
- GRU 기반 순환 신경망은 융합된 공간적 및 시간적 특징을 처리하여 비디오 시퀀스의 장기적 시간적 의존성을 모델링한다.
- 모델은 특징 융합이 최종 정서 예측 헤드 이전에 이루어지는 방식으로 비디오 시퀀스에서 엔드 투 엔드로 훈련된다.
- 조명 변화에 대한 강건성을 평가하기 위해 프레임에 무작위 감마 보정을 적용하여 실제 세계의 조명 변화를 시뮬레이션한다.
실험 결과
연구 질문
- RQ1미세운동 및 거시운동 특징을 융합하면 비디오 정서 인식 성능이 향상되는가?
- RQ2광학 흐름 대신 위상 차이를 사용할 경우 정서 인식 성능과 강건성이 향상되는가?
- RQ3자극성 예측에 비해 미세운동은 정서의 타당성 예측에 얼마나 기여하는가?
- RQ4제안된 운동 표현 방식은 광학 흐름보다 조명 변화에 더 강건한가?
- RQ5공간 스트림과 시간 스트림은 정서 인식에 상호 보완적인 정보를 제공하는가?
주요 결과
- MIMAMO Net은 OMG-Emotion 및 Aff-Wild 데이터셋 모두에서 기존 방법을 능가하는 최신 기술 수준의 성능을 달성한다.
- Aff-Wild 데이터셋에서 공간 스트림과 시간 스트림을 융합했을 때 자극성 예측 성능은 26.8% 향상되었고, 타당성 예측은 12.8% 향상되어 운동 정보가 자극성 예측에 더 관련이 깊다는 것을 확인한다.
- 시간 스트림에 위상 차이를 입력으로 사용할 경우 광학 흐름(0.580 vs. 0.578 F1 점수, Aff-Wild에서 자극성 예측) 및 원시 위상 이미지보다 뛰어난 성능을 기록한다.
- 위상 차이를 사용하는 모델는 조명 변화 하에서도 안정적인 성능을 유지하지만, 광학 흐름 기반 모델은 감마 보정의 변동성이 증가함에 따라 성능이 크게 떨어진다.
- 공간 스트림은 시간 스트림보다 정서 인식에 더 큰 기여를 하지만, 두 스트림을 융합할 경우 성능 향상이 가장 크게 나타나며 특히 자극성 예측에서 두드러진다.
- 위상 차이를 사용함으로써 광학 흐름 대비 약 10배의 계산 비용 절감이 이루어지며, 밝기 변화에 대한 강건성도 향상된다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.