Skip to main content
QUICK REVIEW

[论文解读] A Hybrid Approach to Video Source Identification

Massimo Iuliani, Marco Fontani|arXiv (Cornell University)|May 4, 2017
Digital Media Forensic Detection参考文献 9被引用 8
一句话总结

本文提出了一种混合视频源识别方法,利用从静态图像中提取的传感器图案噪声(SPN)来识别视频源,包括来自相机内数字防抖智能手机的视频。该方法在非防抖视频上实现了最先进性能,并通过在18款现代设备上建模图像与视频SPN之间的几何关系,实现了对原本难以处理的防抖视频的可靠源识别。

ABSTRACT

Multimedia Forensics allows to determine whether videos or images have been captured with the same device, and thus, eventually, by the same person. Currently, the most promising technology to achieve this task, exploits the unique traces left by the camera sensor into the visual content. Anyway, image and video source identification are still treated separately from one another. This approach is limited and anachronistic if we consider that most of the visual media are today acquired using smartphones, that capture both images and videos. In this paper we overcome this limitation by exploring a new approach that allows to synergistically exploit images and videos to study the device from which they both come. Indeed, we prove it is possible to identify the source of a digital video by exploiting a reference sensor pattern noise generated from still images taken by the same device of the query video. The proposed method provides comparable or even better performance, when compared to the current video identification strategies, where a reference pattern is estimated from video frames. We also show how this strategy can be effective even in case of in-camera digitally stabilized videos, where a non-stabilized reference is not available, by solving some state-of-the-art limitations. We explore a possible direct application of this result, that is social media profile linking, i.e. discovering relationships between two or more social media profiles by comparing the visual contents - images or videos - shared therein.

研究动机与目标

  • 解决从智能手机视频中识别源的问题,特别是那些具有相机内数字防抖功能的视频,此类视频无法直接从视频中提取SPN。
  • 克服传统方法依赖设备特定视频训练数据的限制,使用静态图像作为视频指纹的代理。
  • 在包括防抖与非防抖设备在内的多种智能手机之间,建立图像与视频SPN之间的几何模型。
  • 通过统一的指纹识别方法,实现在不同社交平台(如Facebook图片与YouTube视频)之间的跨平台源识别。
  • 开发并发布一个大规模数据集,包含来自18台设备的339段视频和5,289张图像,以支持多媒体取证领域的未来研究。

提出的方法

  • 从同一设备拍摄的高质量静态图像中估计传感器图案噪声(SPN)指纹,将其视为视频源指纹的代理。
  • 使用归一化互相关建模图像与视频SPN之间的几何变换(缩放、裁剪、旋转),以对齐跨模态的指纹。
  • 对缩放和旋转参数进行暴力搜索,以匹配图像衍生的SPN与视频SPN,使用峰值归一化互相关作为匹配度量。
  • 通过基于阈值的融合策略(聚合阈值τ)对多个视频帧的匹配得分进行聚合,以提高鲁棒性并减少误报。
  • 使用零填充的归一化互相关计算SPN块之间的相似性,ρpeak(𝐗, 𝐘)作为主要匹配得分。
  • 利用智能手机相机在图像和视频中使用相同传感器的事实,即使在防抖条件下也能实现在模态间的SPN传递。

实验结果

研究问题

  • RQ1是否可以可靠地使用从静态图像中提取的SPN来识别视频源,尤其是当视频本身无法提供可用的非防抖SPN时?
  • RQ2在包括具有相机内数字防抖功能的现代智能手机在内的设备中,图像与视频SPN之间存在何种几何关系?
  • RQ3该混合方法在不同社交平台(如Facebook与YouTube)之间关联图像与视频内容时的效率如何,这些平台的画质水平各不相同?
  • RQ4当使用低质量图像参考时,实现可靠源识别所需的最少视频帧数是多少?
  • RQ5所提出的方法是否能在传统SPN提取失败的相机内数字防抖视频中实现高精度的源识别?

主要发现

  • 当使用500帧从YouTube视频中估计视频指纹时,该混合方法的AUC达到0.88,性能与传统基于视频的SPN方法相当或更优。
  • 对于用作参考的低质量Facebook图像,使用1,000帧视频可实现AUC为0.86,表明更高的帧数可补偿参考质量的下降。
  • 在最优聚合阈值τ = 38时,该方法在识别相机内数字防抖YouTube视频源方面实现了87.3%的整体准确率。
  • 该方法成功识别了包括最新款苹果设备在内的防抖视频源,而直接从视频中提取SPN在这些设备上不可行。
  • 当视频帧数少于300帧时,性能显著下降,表明存在可靠的指纹估计所必需的最小帧数。
  • 本研究揭示了在18款现代智能手机(包括具有数字防抖功能的设备)中,图像与视频SPN之间存在一致的几何关系,从而实现了可靠的对齐与匹配。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。