Skip to main content
QUICK REVIEW

[论文解读] Aggressive actions and anger detection from multiple modalities using Kinect

Amol Patwardhan, Gerald M. Knapp|arXiv (Cornell University)|Jul 4, 2016
Anomaly Detection Techniques and Applications被引用 10
一句话总结

本文提出了一种基于Kinect的多模态系统,通过融合面部表情、头部与身体动作、手势以及语音信息,实现实时检测愤怒和攻击性行为。结合基于规则的行为特征与支持向量机(SVM),该方法相较单模态方法将愤怒检测的精确率提高了15.2%,召回率提高了11.7%,显著提升了监狱和酒吧等高风险环境中监控系统的响应能力。

ABSTRACT

Prison facilities, mental correctional institutions, sports bars and places of public protest are prone to sudden violence and conflicts. Surveillance systems play an important role in mitigation of hostile behavior and improvement of security by detecting such provocative and aggressive activities. This research proposed using automatic aggressive behavior and anger detection to improve the effectiveness of the surveillance systems. An emotion and aggression aware component will make the surveillance system highly responsive and capable of alerting the security guards in real time. This research proposed facial expression, head, hand and body movement and speech tracking for detecting anger and aggressive actions. Recognition was achieved using support vector machines and rule based features. The multimodal affect recognition precision rate for anger improved by 15.2% and recall rate improved by 11.7% when behavioral rule based features were used in aggressive action detection.

研究动机与目标

  • 提升在监狱、精神卫生机构和酒吧等高风险公共场所中对攻击性行为和愤怒情绪的实时检测能力。
  • 开发一种多模态情感识别系统,整合来自Kinect的视觉与音频线索,以增强行为分析能力。
  • 评估基于规则的行为特征在提升愤怒与攻击行为检测性能方面的有效性。
  • 使监控系统能够通过识别敌意的早期迹象,主动向安保人员发出警报。
  • 证明结合多种模态可显著提高检测准确率,优于单模态方法。

提出的方法

  • 系统使用微软Kinect捕获多模态数据,包括RGB视频、深度图和音频流。
  • 通过面部关键点检测与情绪分类模型分析面部表情。
  • 利用Kinect传感器提供的骨骼关节点数据追踪头部、手部和身体动作。
  • 提取语音特征(如音高、能量和语速)以检测愤怒的语音表现。
  • 设计基于规则的特征,用于建模攻击性行为模式,例如快速手势或突然动作。
  • 使用支持向量机(SVM)对融合的多模态特征进行训练,以分类愤怒与攻击性行为。

实验结果

研究问题

  • RQ1多模态融合面部、动作与语音信号是否能超越单模态方法,提升愤怒与攻击行为的检测效果?
  • RQ2在真实监控场景中,基于规则的行为特征在多大程度上提升了检测性能?
  • RQ3集成Kinect获取的数据如何提升愤怒检测系统的精确率与召回率?
  • RQ4能否利用消费级深度与动作传感器可靠地实现实时攻击行为检测?
  • RQ5各模态(面部表情、身体动作或语音)对整体检测准确率的相对贡献如何?

主要发现

  • 与基线模型相比,引入基于规则的行为特征使愤怒检测的精确率提高了15.2%。
  • 当将基于规则的特征整合进多模态系统时,愤怒检测的召回率提升了11.7%。
  • 面部、动作与语音信号的多模态融合在精确率与召回率上均优于单模态检测方法。
  • 该系统在监狱和酒吧等高风险环境中的实时部署具有可行性。
  • 面部表情与身体动作特征对检测性能的贡献最为显著。
  • 当使用融合的多模态特征进行训练时,支持向量机分类器实现了稳健的分类准确率。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。