Skip to main content
QUICK REVIEW

[论文解读] A Survey for Deep RGBT Tracking

Zhangyong Tang, Tianyang Xu|arXiv (Cornell University)|Jan 23, 2022
Video Surveillance and Tracking Methods被引用 10
一句话总结

本综述全面概述了基于深度神经网络的RGBT跟踪方法,对比了MDNet和Siamese架构在主要基准测试中的表现。研究发现,DMCNet在RGBT210上表现最佳,CMPP在LasHeR上表现最佳,同时建议通过引入端到端架构(如Transformer)和改进的特征融合机制,以提升实时性能与鲁棒性。

ABSTRACT

Visual object tracking with the visible (RGB) and thermal infrared (TIR) electromagnetic waves, shorted in RGBT tracking, recently draws increasing attention in the tracking community. Considering the rapid development of deep learning, a survey for the recent deep neural network based RGBT trackers is presented in this paper. Firstly, we give brief introduction for the RGBT trackers concluded into this category. Then, a comparison among the existing RGBT trackers on several challenging benchmarks is given statistically. Specifically, MDNet and Siamese architectures are the two mainstream frameworks in the RGBT community, especially the former. Trackers based on MDNet achieve higher performance while Siamese-based trackers satisfy the real-time requirement. In summary, since the large-scale dataset LasHeR is published, the integration of end-to-end framework, e.g., Siamese and Transformer, should be further considered to fulfil the real-time as well as more robust performance. Furthermore, the mathematical meaning should be more considered during designing the network. This survey can be treated as a look-up-table for researchers who are concerned about RGBT tracking.

研究动机与目标

  • 提供对近期基于深度神经网络的RGBT跟踪器的系统性综述。
  • 在GTOT、RGBT210、RGBT234、LasHeR以及VOT-RGBT2019/2020等关键基准上,比较现有RGBT跟踪器的性能。
  • 识别主流架构(尤其是MDNet和Siamese网络)的优势与局限性。
  • 突出尚未充分探索的研究方向,包括Transformer和时序建模在RGBT跟踪中的潜力。
  • 通过识别数学严谨性与特征融合机制解释方面的空白,为未来研究提供指导。

提出的方法

  • 将RGBT跟踪器分为三类:基于MDNet的、基于Siamese的以及其他类型。
  • 回顾并对比多模态特征融合中使用的机制,包括注意力模块、自适应融合与门控网络。
  • 在六个基准上评估跟踪器:GTOT、RGBT210、RGBT234、LasHeR、VOT-RGBT2019与VOT-RGBT2020。
  • 采用标准指标分析性能:精确度(↑)、成功率(↑)、准确度(↑)、鲁棒性(↑)与EAO(↑)。
  • 突出DMCNet中模态特定适配器、质量感知融合与互引导注意力等架构创新。
  • 强调端到端训练框架与网络设计中更深层次理论基础的必要性。

实验结果

研究问题

  • RQ1哪些深度学习架构主导了RGBT跟踪领域,它们在性能上如何比较?
  • RQ2基于MDNet与基于Siamese的跟踪器在精度、鲁棒性与实时能力方面有何差异?
  • RQ3在RGBT210、LasHeR与VOT-RGBT2019/2020等主要基准上,表现最佳的RGBT跟踪器是哪些?
  • RQ4为何当前RGBT跟踪方法中,Transformer与端到端学习框架的整合仍处于探索不足状态?
  • RQ5当前特征融合机制的关键局限是什么?数学原理如何能改进其设计?

主要发现

  • 在RGBT210基准上,DMCNet实现了最高的精确度(0.797)与成功率(0.555)。
  • 在RGBT234数据集上,DMCNet同样领先,精确度为0.839,成功率达0.593。
  • 在LasHeR数据集上,APFNet表现最佳,精确度为0.905,成功率达0.739。
  • 在VOT-RGBT2019上,mfDiMP取得最高的EAO得分(0.3879),而DFAT在VOT-RGBT2020挑战中胜出。
  • 在GTOT上,CMPP实现了最高的精确度(0.926),而ADRNet在该基准上取得了最佳成功率(0.739)。
  • 综述指出,尽管基于MDNet的跟踪器精度更高,但基于Siamese的跟踪器更符合实时性要求。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。