Skip to main content
QUICK REVIEW

[논문 리뷰] Audio-Visual Sentiment Analysis for Learning Emotional Arcs in Movies

Eric Chu, Deb Roy|arXiv (Cornell University)|2017. 12. 08.
Media Influence and Health참고 문헌 17인용 수 8
한 줄 요약

이 논문은 영화의 청각적 및 시각적 감정 분석을 위해 딥 컨volution 네트워크를 활용하는 다중모달 기계학습 프레임워크를 제안하며, 관객의 참여도를 예측하는 감정 곡선을 구성한다. 특정 감정 곡선 형태—특히 정점에서 끝나는('끝내기' 형태)—가 온라인 영상 콘텐츠에서 댓글 수가 높은 데 있어 통계적으로 유의미한 예측 요소임을 입증한다.

ABSTRACT

Stories can have tremendous power -- not only useful for entertainment, they can activate our interests and mobilize our actions. The degree to which a story resonates with its audience may be in part reflected in the emotional journey it takes the audience upon. In this paper, we use machine learning methods to construct emotional arcs in movies, calculate families of arcs, and demonstrate the ability for certain arcs to predict audience engagement. The system is applied to Hollywood films and high quality shorts found on the web. We begin by using deep convolutional neural networks for audio and visual sentiment analysis. These models are trained on both new and existing large-scale datasets, after which they can be used to compute separate audio and visual emotional arcs. We then crowdsource annotations for 30-second video clips extracted from highs and lows in the arcs in order to assess the micro-level precision of the system, with precision measured in terms of agreement in polarity between the system's predictions and annotators' ratings. These annotations are also used to combine the audio and visual predictions. Next, we look at macro-level characterizations of movies by investigating whether there exist `universal shapes' of emotional arcs. In particular, we develop a clustering approach to discover distinct classes of emotional arcs. Finally, we show on a sample corpus of short web videos that certain emotional arcs are statistically significant predictors of the number of comments a video receives. These results suggest that the emotional arcs learned by our approach successfully represent macroscopic aspects of a video story that drive audience engagement. Such machine understanding could be used to predict audience reactions to video stories, ultimately improving our ability as storytellers to communicate with each other.

연구 동기 및 목표

  • 영화의 감정 곡선을 오디오 및 시각적 감정 분석을 통해 모델링하여 서사적 영향력에 대한 이해를 향상시키기.
  • 30초 영상 클립에서 인간 애너테이션된 감정 정점과 골절을 기준으로 모델 출력과의 일치도를 비교하여 마이크로 수준 감정 예측 정밀도 평가하기.
  • 시각적 감정 경로의 클러스터링을 통해 보편적인 마크로 수준 감정 곡선 유형을 발견하기.
  • 온라인 영상 플랫폼의 댓글 수를 지표로 삼아 특정 감정 곡선 형태가 관객 참여도를 예측하는지 조사하기.
  • 향후 다중모달 감정 모델링 연구를 위해 오디오 기능과 감정 애너테이션된 영상 클립의 공개 데이터셋 제공하기.

제안 방법

  • 대규모 오디오 및 시각적 감정 데이터셋을 기반으로 딥 컨volution 네트워크를 훈련하여 프레임 수준 및 오디오 스니펫 수준의 감정 점수를 추출하기.
  • 감정 곡선의 형태 유사도를 비교하기 위해 동적 시간 왜곡(DTW)과 Keogh 하한 경계를 적용하기.
  • 실루엣 기반 엣지 방법을 사용한 k-medoids 클러스터링을 통해 시각적 감정 경로에서 구분되는 감정 곡선 유형 가족 식별하기.
  • 가중치 모델을 사용해 오디오 및 시각적 감정 예측을 통합하고, 감정 극성 예측 정확도를 인간 애너테이션 일치도를 기준으로 평가하기.
  • 10% 영상 세그먼트에 걸쳐 양방향 개념 분류기의 2단계 이전 레이어 활성화를 평균화하여 영화 임베딩을 구성한 후, 10×1 특성 벡터로 차원 축소하기.
  • 메타데이터와 범주형 클러스터 할당(감정 곡선 유형 가족)을 사용해 Vimeo 단편 영상의 댓글 수를 예측하기 위한 회귀 분석 수행하기.
Figure 1 : Overview
Figure 1 : Overview

실험 결과

연구 질문

  • RQ1인간 애너테이터의 검증을 통해 오디오-시각적 감정 분석이 30초 영상 클립의 마이크로 수준 감정 정점과 골절을 정확하게 모델링할 수 있는가?
  • RQ2다양한 영화와 단편 영상 간에 구분되는 보편적인 감정 곡선 유형 가족이 존재하는가? 만약 존재한다면, 어떻게 발견할 수 있는가?
  • RQ3일부 특정 감정 곡선 형태가 온라인 댓글 수를 지표로 삼아 관객 참여도를 통계적으로 예측하는가?
  • RQ4오디오 및 시각적 감정 특징의 통합이 감정 곡선 모델링의 정확도와 해석 가능성에 어떤 영향을 미치는가?
  • RQ5감정 곡선 유형 가족은 영화 서사 기법이나 장르와 어느 정도 관련이 있는가?

주요 결과

  • 오디오-시각적 감정 통합 모델은 30초 영상 클립에서 감정 정점과 골절의 극성 예측 정밀도가 0.894에 도달했으며, 인간 애너테이터와의 일치도가 높았다.
  • k-medoids 클러스터링 접근법은 단편 영화 코퍼스에서 다섯 가지 구분되는 감정 곡선 유형 가족을 식별했으며, 'Icarus' 형태(상승-하강)와 'end-with-a-bang' 곡선이 참여도에 대해 통계적으로 유의미한 예측 요소로 나타났다.
  • k=8 클러스터링에서 'end-with-a-bang' 곡선 유형 가족은 댓글 수 예측에 대해 통계적으로 유의미한 예측 요소였으며(p < 0.05), 관객 참여도와 정적 상관관계를 보였다.
  • 'Icarus' 곡선(상승-하강)과 정점에서 끝나는 다른 두 가지 곡선 유형(정점 이전에 상승 또는 수평선)도 여러 k 값에서 높은 댓글 수를 예측하는 데 유의미한 예측 요소로 나타났다.
  • 시각적 감정에서 유도된 영화 임베딩은 장르별로 명확한 클러스터링을 보였으며, 로맨스, 어드벤처, 판타지, 애니메이션 영화가 각각 독립된 그룹을 형성했다.
  • 오디오-시각 융합의 사용은 데이터 희소성과 작은 지표 커버리지로 인해, 시각적 단독 곡선에 비해 덜 뚜렷한 클러스터를 초래했다.
Figure 2 : Effect of smoothing values for arcs: no smoothing versus $w=0.1*n$
Figure 2 : Effect of smoothing values for arcs: no smoothing versus $w=0.1*n$

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.