Skip to main content
QUICK REVIEW

[论文解读] Actions in the Eye: Dynamic Gaze Datasets and Learnt Saliency Models for Visual Recognition

Stefan Mathe, Cristian Sminchisescu|arXiv (Cornell University)|Dec 29, 2013
Visual Attention and Saliency Detection参考文献 42被引用 4
一句话总结

本文介绍了在好莱坞-2和UCF Sports视频动作识别任务中,通过人类眼动追踪获得的大规模动态注视数据集,支持训练能够预测人类注视点的显性注意力模型。通过将这些显性注意力模型整合到端到端可训练的计算机视觉系统中,作者实现了最先进的识别性能,证明人类类似的注意力模式可显著提升视觉动作识别的准确率。

ABSTRACT

Systems based on bag-of-words models from image features collected at maxima of sparse interest point operators have been used successfully for both computer visual object and action recognition tasks. While the sparse, interest-point based approach to recognition is not inconsistent with visual processing in biological systems that operate in `saccade and fixate' regimes, the methodology and emphasis in the human and the computer vision communities remains sharply distinct. Here, we make three contributions aiming to bridge this gap. First, we complement existing state-of-the art large scale dynamic computer vision annotated datasets like Hollywood-2 and UCF Sports with human eye movements collected under the ecological constraints of the visual action recognition task. To our knowledge these are the first large human eye tracking datasets to be collected and made publicly available for video, vision.imar.ro/eyetracking (497,107 frames, each viewed by 16 subjects), unique in terms of their (a) large scale and computer vision relevance, (b) dynamic, video stimuli, (c) task control, as opposed to free-viewing. Second, we introduce novel sequential consistency and alignment measures, which underline the remarkable stability of patterns of visual search among subjects. Third, we leverage the significant amount of collected data in order to pursue studies and build automatic, end-to-end trainable computer vision systems based on human eye movements. Our studies not only shed light on the differences between computer vision spatio-temporal interest point image sampling strategies and the human fixations, as well as their impact for visual recognition performance, but also demonstrate that human fixations can be accurately predicted, and when used in an end-to-end automatic system, leveraging some of the advanced computer vision practice, can lead to state of the art results.

研究动机与目标

  • 通过在视频动作识别任务中收集大规模、任务控制的人眼运动数据,弥合人类视觉注意力与计算机视觉之间的差距。
  • 提出新颖的一致性与对齐度量方法,用于分析不同受试者和视频中人类注视的空间与时序模式。
  • 构建端到端可训练的计算机视觉系统,利用预测的人类显性注意力提升动作识别性能。
  • 评估人类注视模式对识别准确率的影响,并与传统计算机视觉采样策略进行比较。

提出的方法

  • 在受控的动作识别任务下,从16名受试者处收集了来自好莱坞-2和UCF Sports视频数据集的497,107帧的动态眼动追踪数据。
  • 提出时序一致性与对齐度量方法,用于量化不同受试者之间注视模式的空间与时间稳定性。
  • 利用基于显性注意力的采样策略(包括预测和真实显性注意力图)训练端到端的视觉识别流水线。
  • 采用词袋视觉词(bag-of-visual-words)与二阶池化框架,结合Harris角点、均匀采样与显性注意力驱动采样方法的描述符。
  • 使用描述符的协方差矩阵,并结合矩阵对数与幂次缩放,实现在无需额外核函数的情况下的高效非线性特征编码。
  • 采用留一法交叉验证并结合10组随机种子评估结果的方差与鲁棒性。

实验结果

研究问题

  • RQ1在动态视频动作识别任务中,人类注视模式在不同受试者之间是否具有一致性?
  • RQ2任务约束在多大程度上影响了视频中人类眼动的空间与时序结构?
  • RQ3基于人类眼动追踪数据训练的显性注意力模型能否准确预测视频刺激中的注视位置?
  • RQ4与传统兴趣点或均匀采样相比,基于显性注意力的采样在识别准确率方面表现如何?
  • RQ5使用预测的人类显性注意力的端到端可训练系统能否在视觉动作识别中实现最先进性能?

主要发现

  • 即使在任务约束下,人类注视模式在空间与时序上均表现出高度一致性,表明视觉搜索行为具有稳定性。
  • 基于人类注视数据训练的显性注意力模型在预测视频刺激中的注视位置方面表现出高准确性。
  • 在UCF Sports数据集上,采用基于预测显性注意力的采样策略在词袋视觉词框架中将识别准确率提升至87.5%,优于基线方法。
  • 在10组随机种子下,识别准确率的标准差低于0.8%,表明所提流水线具有低方差与高鲁棒性。
  • 基于显性注意力的采样优于传统方法,如Harris角点(84.3%)与均匀采样(83.9%),证明了人类注意力先验的价值。
  • 结合显性注意力采样的二阶池化方法实现了最先进性能,证实人类类似的注意力机制可增强计算机视觉系统。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。