Skip to main content
QUICK REVIEW

[论文解读] Large Margin Structured Convolution Operator for Thermal Infrared Object Tracking

Peng Gao, Yipeng Ma|arXiv (Cornell University)|Apr 19, 2018
Video Surveillance and Tracking Methods参考文献 33被引用 5
一句话总结

本文提出了一种新型热成像(TIR)目标跟踪方法——大 margin 结构化卷积算子(LMSCO),该方法将结构化输出支持向量机(SOSVM)的判别能力与判别相关滤波器(DCF)的高效性相结合。通过利用预训练网络提取的深度外观与运动特征,应用空间正则化和隐式插值以生成连续的特征图,并采用在线协同优化策略,LMSCO 在 VOT-TIR2015 和 VOT-TIR2016 基准上实现了最先进水平的准确率与鲁棒性,同时具备高帧率性能。

ABSTRACT

Compared with visible object tracking, thermal infrared (TIR) object tracking can track an arbitrary target in total darkness since it cannot be influenced by illumination variations. However, there are many unwanted attributes that constrain the potentials of TIR tracking, such as the absence of visual color patterns and low resolutions. Recently, structured output support vector machine (SOSVM) and discriminative correlation filter (DCF) have been successfully applied to visible object tracking, respectively. Motivated by these, in this paper, we propose a large margin structured convolution operator (LMSCO) to achieve efficient TIR object tracking. To improve the tracking performance, we employ the spatial regularization and implicit interpolation to obtain continuous deep feature maps, including deep appearance features and deep motion features, of the TIR targets. Finally, a collaborative optimization strategy is exploited to significantly update the operators. Our approach not only inherits the advantage of the strong discriminative capability of SOSVM but also achieves accurate and robust tracking with higher-dimensional features and more dense samples. To the best of our knowledge, we are the first to incorporate the advantages of DCF and SOSVM for TIR object tracking. Comprehensive evaluations on two thermal infrared tracking benchmarks, i.e. VOT-TIR2015 and VOT-TIR2016, clearly demonstrate that our LMSCO tracker achieves impressive results and outperforms most state-of-the-art trackers in terms of accuracy and robustness with sufficient frame rate.

研究动机与目标

  • 通过利用深度特征与鲁棒学习框架,解决热成像目标跟踪中缺乏颜色模式与低分辨率等局限性。
  • 通过整合高维深度外观与运动特征,提升热成像中的跟踪性能。
  • 通过将 SOSVM 与高效 DCF 框架融合,克服 SOSVM 的计算瓶颈,实现实时推理。
  • 开发一种协同优化策略,同时提升跟踪精度与速度。

提出的方法

  • 提出一种大 margin 结构化卷积算子(LMSCO),将 SOSVM 的判别能力与 DCF 的计算效率相结合。
  • 采用空间正则化与隐式插值,从 TIR 序列生成连续且高质量的深度特征图。
  • 将预训练于可见光图像的卷积神经网络(CNN)迁移至 TIR 跟踪任务,以提取深度外观与运动特征,实现更丰富的表征。
  • 在 SOSVM 框架内引入循环矩阵与相关滤波器,实现基于密集采样的快速在线训练。
  • 提出一种在线协同优化策略,以在推理过程中高效更新跟踪算子。
  • 结合深度外观与运动特征,提升在低纹理、低分辨率 TIR 条件下的特征多样性与鲁棒性。

实验结果

研究问题

  • RQ1SOSVM 与 DCF 框架的融合是否能提升热成像目标跟踪中的精度与速度?
  • RQ2从预训练于可见光域的网络中提取的深度外观与运动特征,在 TIR 跟踪中是否有效?
  • RQ3空间正则化与隐式插值是否能提升低分辨率 TIR 序列中深度特征图的质量?
  • RQ4在线协同优化策略是否能显著提升 TIR 跟踪中的性能与推理速度?
  • RQ5所提出方法是否能在标准 TIR 基准上超越最先进跟踪器,在精度、鲁棒性与帧率方面表现更优?

主要发现

  • LMSCO 在 VOT-TIR2016 基准上取得 42.8% 的 EAO 得分,优于先前最先进跟踪器 SRDCF(36.4%),绝对提升 6.4%。
  • 该跟踪器在 VOT-TIR2016 的所有评估指标中均表现最佳,包括均值、合并与加权均值 AR 图。
  • LMSCO 在所有评估协议下均显著优于九种最先进跟踪器,包括 MDNet、SRDCF、Staple 与 GGT2。
  • 深度外观与运动特征的融合显著提升了跟踪性能,优于单一特征基线与手工设计特征。
  • 所提出的协同优化策略在保持高精度与鲁棒性的同时,实现了高帧率的实时跟踪。
  • 图 1 的可视化对比显示,LMSCO 在 'hiding' 与 'quadrocopter' 等挑战性序列中能始终保持更紧密且一致的边界框。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。