Skip to main content
QUICK REVIEW

[论文解读] Non-Volume Preserving-based Feature Fusion Approach to Group-Level Expression Recognition on Crowd Videos.

Kha Gia Quach, Ngan Le|arXiv (Cornell University)|Nov 28, 2018
Emotion and Mood Recognition参考文献 25被引用 6
一句话总结

本文提出了一种非体积保持融合(NVPF)框架,用于人群视频中的群体级表情识别,通过深度特征级融合来建模空间与时间关系。该方法在新的群体级人群视频表情数据集(GECV)上实现了最先进性能,展示了在个体、群体和帧级别识别情绪的鲁棒性。

ABSTRACT

Group-level emotion recognition (ER) is a growing research area as the demands for assessing crowds of all sizes is becoming an interest in both the security arena as well as social media. This work extends the earlier ER investigations, which focused on either group-level ER on single images or within a video, by fully investigating group-level expression recognition on crowd videos. In this paper, we propose an effective deep feature level fusion mechanism to model the spatial-temporal information in the crowd videos. In our approach, the fusing process is performed on deep feature domain by a generative probabilistic model, Non-Volume Preserving Fusion (NVPF), that models spatial information relationship. Furthermore, we extend our proposed spatial NVPF approach to spatial-temporal NVPF approach to learn the temporal information between frames. In order to demonstrate the robustness and effectiveness of each component in the proposed approach, three experiments were conducted: (i) evaluation on AffectNet database to benchmark the proposed EmoNet for recognizing facial expression; (ii) evaluation on EmotiW2018 to benchmark the proposed deep feature level fusion mechanism NVPF; and, (iii) examine the proposed TNVPF on an innovative Group-level Emotion on Crowd Videos (GECV) dataset composed of 627 videos collected from publicly available sources. GECV dataset is a collection of videos containing crowds of people. Each video is labeled with emotion categories at three levels: individual faces, group of people and the entire video frame.

研究动机与目标

  • 为解决现有研究在完整人群视频中群体级表情识别的空白,超越单幅图像或孤立帧的局限。
  • 开发一种深度特征级融合机制,有效建模人群视频中的空间与时间关系。
  • 创建并发布一个新的基准数据集GECV,包含对个体、群体和整个帧的多层次情绪标注。
  • 在真实世界的人群视频数据上评估所提出的NVPF和时序NVPF(TNVPF)方法,以提升情绪识别性能。
  • 通过在多个数据集上的消融实验与对比实验,证明融合机制的有效性。

提出的方法

  • 提出非体积保持融合(NVPF),一种生成式概率模型,可在保留人群视频中空间关系的同时融合深度特征。
  • 将空间NVPF扩展为时空NVPF(TNVPF),以建模视频帧间的时间动态。
  • 在深度特征域中应用NVPF机制,融合来自多帧和面部区域的表征。
  • 采用概率框架建模特征交互,无需假设体积保持,从而实现灵活且富有表现力的融合。
  • 在GECV数据集中采用多层次标注方案,对个体、群体和帧级别的情绪进行标注,用于训练与评估。
  • 在三个基准数据集上验证该方法:用于面部表情识别的AffectNet、用于特征融合的EmotiW2018,以及用于群体级识别的新GECV数据集。

实验结果

研究问题

  • RQ1所提出的NVPF机制在人群视频中群体级表情识别的深度特征融合方面效果如何?
  • RQ2时空扩展(TNVPF)是否通过建模帧间的时间动态提升性能?
  • RQ3在具有多层次情绪标注的真实世界人群视频数据上,该方法与现有方法相比表现如何?
  • RQ4空间建模与时间建模两个组件对整体识别性能的贡献分别是什么?
  • RQ5在多样化、公开获取的人群视频上评估时,该方法的鲁棒性如何?

主要发现

  • 所提出的NVPF方法在EmotiW2018基准上实现了最先进性能,证明其在建模复杂特征交互方面的有效性。
  • TNVPF扩展显著提升了识别准确率,能够捕捉人群视频中帧间的时间依赖性。
  • 在新引入的GECV数据集中,该方法在个体、群体和帧三个情绪标注层级上均表现出色。
  • 消融实验确认,空间建模与时间建模两个组件均对最终识别准确率有显著贡献。
  • 该方法在多样化视频源上表现出鲁棒性,表明其在真实世界人群视频数据上的泛化能力。
  • GECV数据集为未来群体级表情识别研究提供了宝贵的基准,包含627个视频及多层次标注。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。