[논문 리뷰] MOSI: Multimodal Corpus of Sentiment Intensity and Subjectivity Analysis in Online Opinion Videos
논문은 온라인 비디오에서 감정 강도와 주관성에 대한 의견 수준의 다중모달 말뭉치 MOSI를 소개합니다. 프레임별 시각 특징 및 밀리초 단위의 오디오 특징, 그리고 베이스라인 및 다중모달 융합 모델을 포함합니다.
People are sharing their opinions, stories and reviews through online video sharing websites every day. Studying sentiment and subjectivity in these opinion videos is experiencing a growing attention from academia and industry. While sentiment analysis has been successful for text, it is an understudied research question for videos and multimedia content. The biggest setbacks for studies in this direction are lack of a proper dataset, methodology, baselines and statistical analysis of how information from different modality sources relate to each other. This paper introduces to the scientific community the first opinion-level annotated corpus of sentiment and subjectivity analysis in online videos called Multimodal Opinion-level Sentiment Intensity dataset (MOSI). The dataset is rigorously annotated with labels for subjectivity, sentiment intensity, per-frame and per-opinion annotated visual features, and per-milliseconds annotated audio features. Furthermore, we present baselines for future studies in this direction as well as a new multimodal fusion approach that jointly models spoken words and visual gestures.
연구 동기 및 목표
- 온라인 비디오의 감정 및 주관성에 대한 적절한 다중모달 데이터셋 부재를 동기 부여하고 해결한다.
- 시각적, 음향 및 발화 콘텐츠를 포함한 풍부한 모달리티 주석이 있는 의견 수준 주석 말뭉치를 제공한다.
- 비디오 데이터에서 다중모달 감정 분석 및 주관성 탐지의 기준점을 확립한다.
- 발화 단어와 시각 제스처를 공동으로 모델링하는 다중모달 융합 접근법을 제안한다.
제안 방법
- 온라인 비디오의 감정 및 주관성에 대한 최초의 의견 수준 주석 말뭉치로 MOSI를 도입한다.
- 주관성 라벨, 감정 강도, 프레임별 시각 특징, 의견별 주석 및 밀리초 단위 오디오 특징으로 데이터를 주석화한다.
- 다중모달 감정 분석의 향후 연구를 위한 베이스라인 모델을 제공한다.
- 발화 단어와 시각 제스처를 공동으로 모델링하는 새로운 다중모달 융합 접근법을 제안한다.
실험 결과
연구 질문
- RQ1온라인 비디오에서 의견 수준으로 어떻게 감정 강도와 주관성을 효과적으로 주석 달고 측정할 수 있는가?
- RQ2텍스트, 오디오 및 시각 단서를 결합한 비디오 데이터의 다중모달 감정 분석에 적합한 베이스라인은 무엇인가?
- RQ3발화 단어와 시각 제스처를 공동으로 사용하는 융합 모델이 단일 모달 접근법보다 감정 및 주관성 분석을 개선할 수 있는가?
주요 결과
- MOSI는 온라인 비디오에서 의견 수준의 감정 및 주관성에 대해 엄격하게 주석된 말뭉치를 제공한다.
- 데이터세트는 프레임당 시각 특징 및 밀리초 단위 오디오 특징을 포함하여 세밀한 분석을 지원한다.
- 베이스라인 모델과 새로운 다중모달 융합 접근법이 발화 내용과 시각 제스처를 함께 모델링하도록 제안된다.
- 연구는 비디오 데이터에서 다중모달 감정 분석 연구를 위한 기반을 확립한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.