Skip to main content
QUICK REVIEW

[论文解读] Human Attention Detection Using AM-FM Representations

Wenjing Shi|arXiv (Cornell University)|Mar 9, 2022
Infrared Target Detection Methodologies被引用 8
一句话总结

本文提出了一种基于相位的人类注意力检测方法,采用幅度调制-频率调制(AM-FM)模型,在非受控视频环境中检测人脸朝向、后脑勺存在情况以及视线方向。在AOLME数据集的19,387张图像上进行评估,当面向摄像头时,左视和右视检测的准确率分别达到97.1%和95.9%;在后脑勺情况下,左视和右视检测的准确率分别为87.6%和93.3%,表明该方法在无受控成像几何条件下的真实世界环境中具有强大的鲁棒性。

ABSTRACT

Human activity detection from digital videos presents many challenges to the computer vision and image processing communities. Recently, many methods have been developed to detect human activities with varying degree of success. Yet, the general human activity detection problem remains very challenging, especially when the methods need to work 'in the wild' (e.g., without having precise control over the imaging geometry). The thesis explores phase-based solutions for (i) detecting faces, (ii) back of the heads, (iii) joint detection of faces and back of the heads, and (iv) whether the head is looking to the left or the right, using standard video cameras without any control on the imaging geometry. The proposed phase-based approach is based on the development of simple and robust methods that rely on the use of Amplitude Modulation- Frequency Modulation (AM-FM) models. The approach is validated using video frames extracted from the Advancing Out-of-school Learning in Mathematics and Engineering (AOLME) project. The dataset consisted of 13,265 images from ten students looking at the camera, and 6,122 images from five students looking away from the camera. For the students facing the camera, the method was able to correctly classify 97.1% of them looking to the left and 95.9% of them looking to the right. For the students facing the back of the camera, the method was able to correctly classify 87.6% of them looking to the left and 93.3% of them looking to the right. The results indicate that AM-FM based methods hold great promise for analyzing human activity videos.

研究动机与目标

  • 开发一种在无受控成像几何条件下的非受控视频环境中,具备鲁棒性的基于相位的人类注意力检测方法。
  • 解决在真实世界环境中使用普通摄像机检测人类头部朝向和视线方向的挑战。
  • 探索AM-FM模型在同时检测人脸、后脑勺和视线方向方面的可行性。
  • 在从教育环境中学生收集的真实世界数据集上验证该方法。

提出的方法

  • 该方法采用幅度调制-频率调制(AM-FM)模型,从视频帧中提取基于相位的特征,突出显示结构和方向信息。
  • 通过经验模态分解(EMD)和希尔伯特变换提取相位信息,以建模幅度调制和频率调制。
  • 该方法利用局部相位一致性与方向特征,检测头部区域并估计头部姿态。
  • 应用分类流水线以检测头部是否正对前方(正面或后脑勺)以及视线方向(左或右)。
  • 该方法设计简洁且鲁棒,最大限度减少对精确校准或受控光照条件的依赖。
  • 从视频帧中提取特征,并输入至在AOLME数据集标注数据上训练的分类器。

实验结果

研究问题

  • RQ1AM-FM表征能否有效检测非受控视频帧中的人脸和后脑勺?
  • RQ2在无受控成像几何条件下,基于相位的特征能否准确估计视线方向(左或右)?
  • RQ3当应用于具有可变光照和相机角度的真实世界视频数据时,AM-FM模型是否保持鲁棒性?
  • RQ4统一的基于相位的框架能否同时检测多种与注意力相关的线索(人脸、后脑勺、视线)?
  • RQ5在非受控环境中,AM-FM方法的性能与传统方法相比如何?

主要发现

  • 该方法在面向摄像头的学生中,左视检测的分类准确率达到97.1%。
  • 对于面向摄像头的学生,95.9%的右视情况被正确分类。
  • 在检测后脑勺朝向时,87.6%的左视情况被正确识别。
  • 对于后脑勺情况,93.3%的右视情况被正确分类。
  • 总体结果表明,该方法在真实世界、非受控视频环境中表现出强大的鲁棒性和泛化能力。
  • 基于AM-FM的方法在复杂、非受控环境中优于传统的基于强度的方法。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。