Skip to main content
QUICK REVIEW

[论文解读] Efficient Video Summarization Framework using EEG and Eye-tracking Signals

Sai Sukruth Bezugam, Swatilekha Majumdar|arXiv (Cornell University)|Jan 27, 2021
Video Analysis and Summarization参考文献 42被引用 6
一句话总结

本文提出了一种以人为中心的视频摘要框架,利用脑电图(EEG)和眼动追踪信号识别感知显著帧,实现96.5%的视频压缩率,精度为0.98,召回率为0.97——在优于最先进方法的同时,通过生物启发的可解释特征提取显著降低了计算成本。

ABSTRACT

This paper proposes an efficient video summarization framework that will give a gist of the entire video in a few key-frames or video skims. Existing video summarization frameworks are based on algorithms that utilize computer vision low-level feature extraction or high-level domain level extraction. However, being the ultimate user of the summarized video, humans remain the most neglected aspect. Therefore, the proposed paper considers human's role in summarization and introduces human visual attention-based summarization techniques. To understand human attention behavior, we have designed and performed experiments with human participants using electroencephalogram (EEG) and eye-tracking technology. The EEG and eye-tracking data obtained from the experimentation are processed simultaneously and used to segment frames containing useful information from a considerable video volume. Thus, the frame segmentation primarily relies on the cognitive judgments of human beings. Using our approach, a video is summarized by 96.5% while maintaining higher precision and high recall factors. The comparison with the state-of-the-art techniques demonstrates that the proposed approach yields ceiling-level performance with reduced computational cost in summarising the videos.

研究动机与目标

  • 解决现有视频摘要方法忽视人类感知与注意力行为作为核心输入的缺陷。
  • 开发一种计算高效、可解释的视频摘要框架,其基础为人认知信号。
  • 评估结合脑电图(EEG)与眼动追踪数据在识别代表视频核心内容的关键帧方面的有效性。
  • 在压缩率、精度与召回率方面,将所提方法与最先进技术进行基准对比。
  • 通过整合神经生理信号与行为信号,为人类-机器协同视频摘要奠定基础。

提出的方法

  • 通过控制实验,使用脑电图(EEG)与眼动追踪设备,采集参与者在观看视频时的神经反应与视觉注意力响应。
  • 并行处理脑电图(EEG)与眼动追踪信号,检测认知与视觉注意力事件,识别具有高感知显著性的帧。
  • 利用融合的注意力信号,基于人类对相关性的判断,对长视频序列进行分段与关键帧提取。
  • 应用帧选择机制,优先选取脑电图(EEG)与眼动追踪信号响应强烈的部分,表明其信息含量高。
  • 优化摘要处理流程,在最小计算开销下实现高压缩率,避免使用深度学习架构。
  • 通过真实标注数据验证方法,并与传统方法及基于深度学习的摘要基线进行性能对比。

实验结果

研究问题

  • RQ1脑电图(EEG)与眼动追踪信号在识别用于摘要的感知相关视频帧方面有多高效?
  • RQ2与单独使用任一模态相比,结合脑电图(EEG)与眼动追踪信号能带来多大的性能提升?
  • RQ3非深度学习、人机协同的框架是否能在计算成本更低的前提下,实现与最先进方法相当的性能?
  • RQ4所提框架在实现高视频压缩率的同时,能在多大程度上保持高精度与高召回率?
  • RQ5人类感知信号的整合在多大程度上提升了视频摘要系统的可解释性与可解释性?

主要发现

  • 所提框架实现了96.5%的视频压缩率,意味着仅保留原始视频长度的3.5%作为摘要。
  • 该方法的精度约为0.98,表明98%的提取关键帧与视频内容相关。
  • 召回率约为0.97,表明系统成功检索出97%的真实相关帧。
  • 脑电图(EEG)与眼动追踪信号的结合检测准确率达到98.15%,显著优于仅使用脑电图(EEG)(68.7%)或仅使用眼动追踪(15.27%)的方法。
  • 所提方法的F-score最高,证实其在关键帧提取方面整体准确性优于现有框架。
  • 该框架在精度与召回率方面均优于基于深度学习的方法,同时参数量更少,且具备更高的可解释性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。