Skip to main content
QUICK REVIEW

[论文解读] Predicting emotion from music videos: exploring the relative contribution of visual and auditory information to affective responses

Phoebe Chua, Dimos Makris|arXiv (Cornell University)|Feb 19, 2022
Music and Audio Processing被引用 7
一句话总结

本文介绍了 MuVi 数据集,这是一个针对音乐视频在仅音频、仅视频及音视频结合三种条件下进行愉悦度(valence)与唤醒度(arousal)标注的新数据集。本文提出了 PAIR 模型,一种迁移学习架构,利用孤立模态的评分来提升多模态情绪预测性能,发现听觉输入主要驱动唤醒度感知,而两种模态共同影响愉悦度感知。

ABSTRACT

Although media content is increasingly produced, distributed, and consumed in multiple combinations of modalities, how individual modalities contribute to the perceived emotion of a media item remains poorly understood. In this paper we present MusicVideos (MuVi), a novel dataset for affective multimedia content analysis to study how the auditory and visual modalities contribute to the perceived emotion of media. The data were collected by presenting music videos to participants in three conditions: music, visual, and audiovisual. Participants annotated the music videos for valence and arousal over time, as well as the overall emotion conveyed. We present detailed descriptive statistics for key measures in the dataset and the results of feature importance analyses for each condition. Finally, we propose a novel transfer learning architecture to train Predictive models Augmented with Isolated modality Ratings (PAIR) and demonstrate the potential of isolated modality ratings for enhancing multimodal emotion recognition. Our results suggest that perceptions of arousal are influenced primarily by auditory information, while perceptions of valence are more subjective and can be influenced by both visual and auditory information. The dataset is made publicly available.

研究动机与目标

  • 研究听觉与视觉模态在音乐视频中独立与联合作用于感知情绪的方式。
  • 构建一个数据集,捕捉在孤立与组合模态下维度化(愉悦度/唤醒度)与离散情绪的标注。
  • 设计一种新颖的架构 PAIR,利用孤立模态评分作为辅助监督,提升多模态情绪识别性能。
  • 分析性别、音乐训练与歌曲熟悉度等人口统计因素对情绪感知的影响。
  • 探索视觉上下文特征(场景、物体与动作)在时序情绪预测中的实用性。

提出的方法

  • 收集音乐视频,并在三种条件下呈现:音视频(原始)、仅音频(音乐)与仅视频(静音),参与者对愉悦度与唤醒度进行时间序列标注。
  • 收集标注者元数据,包括性别、对歌曲的熟悉度与音乐训练背景,以分析人口统计因素对情绪感知的影响。
  • 使用基于梯度的方法进行特征重要性分析,识别驱动情绪预测的关键音频与视觉特征。
  • 提出 PAIR(Predictive model Augmented with Isolated modality Ratings),一种迁移学习框架,通过联合训练音视频数据与孤立模态评分,提升泛化能力。
  • 使用长短期记忆网络(LSTM)与多模态融合策略训练并评估模型,以预测连续的愉悦度与唤醒度分数。
  • 使用评分者间一致性(ICC)与相关性分析,验证不同模态与条件下情绪标注的一致性。

实验结果

研究问题

  • RQ1孤立的听觉与视觉模态在音乐视频中对感知唤醒度与愉悦度的贡献有何不同?
  • RQ2性别、音乐训练与歌曲熟悉度等人口统计因素在多大程度上影响音乐视频中的情绪感知?
  • RQ3视觉特征中,场景、物体或动作哪一类携带最多的情绪信息?
  • RQ4孤立模态评分能否提升多模态情绪识别模型的性能?
  • RQ5在何种条件下,视觉或听觉主导性会在多模态情绪感知中显现?

主要发现

  • 听觉信息解释了感知唤醒度的大部分方差,表明音乐是音乐视频中唤醒度感知的主要驱动因素。
  • 听觉与视觉模态均对感知愉悦度有贡献,其中视觉内容在塑造主观愉悦度判断中起着重要作用。
  • 视觉模态在愉悦度上表现出最高的评分者间一致性,但在唤醒度上最低,表明在观看静音视频时,愉悦度感知更具共识。
  • 孤立模态评分显示,与仅音乐条件相比,加入视觉内容通常会降低感知的愉悦度与唤醒度。
  • 特征重要性分析表明,与动作相关的视觉特征对情绪识别尤为关键,而基于场景的特征更适合时间建模,而非固定时长的窗口。
  • PAIR 模型通过整合孤立模态评分,在多模态情绪预测中表现出更优性能,证实了辅助模态监督的价值。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。