[Paper Review] MOSI: Multimodal Corpus of Sentiment Intensity and Subjectivity Analysis in Online Opinion Videos
The paper introduces MOSI, the first opinion-level multimodal corpus annotated for sentiment intensity and subjectivity in online videos, with per-frame visual features and per-millisecond audio features, plus baselines and a multimodal fusion model.
People are sharing their opinions, stories and reviews through online video sharing websites every day. Studying sentiment and subjectivity in these opinion videos is experiencing a growing attention from academia and industry. While sentiment analysis has been successful for text, it is an understudied research question for videos and multimedia content. The biggest setbacks for studies in this direction are lack of a proper dataset, methodology, baselines and statistical analysis of how information from different modality sources relate to each other. This paper introduces to the scientific community the first opinion-level annotated corpus of sentiment and subjectivity analysis in online videos called Multimodal Opinion-level Sentiment Intensity dataset (MOSI). The dataset is rigorously annotated with labels for subjectivity, sentiment intensity, per-frame and per-opinion annotated visual features, and per-milliseconds annotated audio features. Furthermore, we present baselines for future studies in this direction as well as a new multimodal fusion approach that jointly models spoken words and visual gestures.
Motivation & Objective
- Motivate and address the lack of proper multimodal datasets for sentiment and subjectivity in online videos.
- Provide an opinion-level annotated corpus with rich modality annotations (visual, audio, and spoken content).
- Establish baselines for multimodal sentiment analysis and subjectivity detection on video data.
- Propose a multimodal fusion approach that jointly models spoken words and visual gestures.
Proposed method
- Introduce MOSI as the first opinion-level annotated corpus for sentiment and subjectivity in online videos.
- Annotate data with subjectivity labels, sentiment intensity, per-frame visual features, per-opinion annotations, and per-millisecond audio features.
- Provide baseline models for future research in multimodal sentiment analysis.
- Propose a new multimodal fusion approach that jointly models spoken words and visual gestures.
Experimental results
Research questions
- RQ1How can sentiment intensity and subjectivity be effectively annotated and measured at the opinion level in online videos?
- RQ2What baselines are suitable for multimodal sentiment analysis on video data combining text, audio, and visual cues?
- RQ3Can a fusion model that jointly uses spoken words and visual gestures improve sentiment and subjectivity analysis over unimodal approaches?
Key findings
- MOSI provides a rigorously annotated corpus for opinion-level sentiment and subjectivity in online videos.
- The dataset includes per-frame visual features and per-millisecond audio features to support fine-grained analysis.
- Baseline models and a new multimodal fusion approach are proposed to jointly model spoken content and visual gestures.
- The study establishes a foundation for multimodal sentiment analysis research on video data.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.