[论文解读] Fusing Audio, Textual and Visual Features for Sentiment Analysis of News Videos
本文提出一种多模态方法,融合音频、文本和视觉特征——如面部情绪强度、语音基频、声音有声度概率、音量以及字幕中的情感得分——以分类新闻视频中的紧张程度(低/高)。在520段巴西和美国新闻视频上评估,该方法达到84%的准确率,显示出在媒体与新闻分析应用中的强大潜力。
This paper presents a novel approach to perform sentiment analysis of news videos, based on the fusion of audio, textual and visual clues extracted from their contents. The proposed approach aims at contributing to the semiodiscoursive study regarding the construction of the ethos (identity) of this media universe, which has become a central part of the modern-day lives of millions of people. To achieve this goal, we apply state-of-the-art computational methods for (1) automatic emotion recognition from facial expressions, (2) extraction of modulations in the participants' speeches and (3) sentiment analysis from the closed caption associated to the videos of interest. More specifically, we compute features, such as, visual intensities of recognized emotions, field sizes of participants, voicing probability, sound loudness, speech fundamental frequencies and the sentiment scores (polarities) from text sentences in the closed caption. Experimental results with a dataset containing 520 annotated news videos from three Brazilian and one American popular TV newscasts show that our approach achieves an accuracy of up to 84% in the sentiments (tension levels) classification task, thus demonstrating its high potential to be used by media analysts in several applications, especially, in the journalistic domain.
研究动机与目标
- 开发一种计算方法,通过整合语言和非语言线索,支持新闻视频的半语 discourse 分析。
- 通过融合音频、文本和视觉模态,弥补现有研究在推断新闻广播中情感紧张方面的空白。
- 在单模态情感分析基础上,提升新闻视频中紧张程度分类的准确性。
- 支持媒体分析师和记者识别情绪化内容,以实现内容摘要、组织与受众参与。
- 验证字幕中情感得分作为检测高紧张新闻片段关键信号的有效性。
提出的方法
- 使用面部表情识别提取视觉特征,量化视频帧中情绪强度(如快乐、愤怒、悲伤)。
- 分析包括语音基频、有声度概率和音量在内的音频特征,以检测表明紧张的语音调制。
- 使用最先进的情感分析技术处理字幕中的文本内容,提取极性得分(正面/负面情感)。
- 在特征级和决策级融合多模态特征,并根据其对紧张程度推断的贡献进行加权。
- 基于视频持续时间内所有多模态特征的加权总和值,将每个新闻视频分类为低紧张或高紧张。
- 通过四位人工标注者的投票机制定义真实标签,评估结果在完整数据集和100%一致子集上进行。
实验结果
研究问题
- RQ1与单模态方法相比,音频、文本和视觉特征的融合是否能提升新闻视频中紧张程度分类的准确性?
- RQ2与视觉和音频线索相比,字幕中的情感得分在识别高紧张新闻片段方面有多有效?
- RQ3所提出的多模态融合方法是否在分类低紧张新闻视频方面优于单一特征集?
- RQ4面部表情和语音调制在多大程度上与新闻广播中感知到的情感紧张相关?
- RQ5自动化紧张程度检测能否支持新闻摘要和媒体分析中的内容优先级排序等实际应用?
主要发现
- 所提出的多模态融合方法在520段新闻视频的数据集上,分类紧张程度(低或高)的准确率达到84%。
- 当限制在人工标注者100%一致的视频时,该方法的准确率提升至85%,证实了在一致标注下的鲁棒性。
- 字幕中的情感得分是识别高紧张新闻的最有效单一特征,在100%一致子集上达到83%的准确率。
- 在分类低紧张新闻方面,该方法优于单一特征(如特征尺寸、情感得分),其中情感得分单独表现较差。
- 多模态特征的融合显著提升了性能,与单模态方法相比,配对t检验在95%置信水平下确认了统计显著性。
- 研究表明,对字幕进行情感分析是多模态新闻视频分析中一种有价值且未被充分利用的信号,尤其在检测情绪化内容方面。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。