Skip to main content
QUICK REVIEW

[논문 리뷰] Exploring Fusion Strategies for Accurate RGBT Visual Object Tracking

Zhangyong Tang, Tianyang Xu|arXiv (Cornell University)|2022. 01. 21.
Advanced Image Fusion Techniques인용 수 6
한 줄 요약

이 논문은 정확한 RGBT 시각적 객체 추적을 위한 새로운 결합 수준 융합 전략인 DFAT을 제안한다. DFAT은 RGB 및 TIR 기여도를 동적으로 가중하고 선형 템플릿 업데이트를 적용한다. RGB 데이터로만 훈련되었지만, VOT-RGBT2020에서 최신 기준 성능을 달성하였으며, 새로운 SOTA EAO 점수 0.4178를 기록하였고, VOT-RGBT2020 도전 대회에서 우승하였다.

ABSTRACT

We address the problem of multi-modal object tracking in video and explore various options of fusing the complementary information conveyed by the visible (RGB) and thermal infrared (TIR) modalities including pixel-level, feature-level and decision-level fusion. Specifically, different from the existing methods, paradigm of image fusion task is heeded for fusion at pixel level. Feature-level fusion is fulfilled by attention mechanism with channels excited optionally. Besides, at decision level, a novel fusion strategy is put forward since an effortless averaging configuration has shown the superiority. The effectiveness of the proposed decision-level fusion strategy owes to a number of innovative contributions, including a dynamic weighting of the RGB and TIR contributions and a linear template update operation. A variant of which produced the winning tracker at the Visual Object Tracking Challenge 2020 (VOT-RGBT2020). The concurrent exploration of innovative pixel- and feature-level fusion strategies highlights the advantages of the proposed decision-level fusion method. Extensive experimental results on three challenging datasets, extit{i.e.}, GTOT, VOT-RGBT2019, and VOT-RGBT2020, demonstrate the effectiveness and robustness of the proposed method, compared to the state-of-the-art approaches. Code will be shared at extcolor{blue}{\emph{https://github.com/Zhangyong-Tang/DFAT}.

연구 동기 및 목표

  • 일조 변화, 가림, 흐림 등의 악조건 하에서 강건한 시각적 객체 추적 문제를 해결한다.
  • 기존 융합 전략(픽셀 수준, 특징 수준, 결합 수준)의 한계를 극복하기 위해 RGBT 추적에서 이들의 상대적 효과성을 비교 분석한다.
  • 쌍방의 RGB-TIR 데이터로 공동 훈련이 필요 없이 RGB 및 열화상(TIR) 모odalities 간의 상보적 정보를 활용하는 결합 수준 융합 방법을 개발한다.
  • 단순히 RGB 데이터로만 훈련된 경량이고 효율적인 트래커를 설계하여 추론 시 TIR 모달리티를 동적 가중치와 템플릿 업데이트를 통해 융합할 수 있도록 한다.
  • 특히 VOT-RGBT2020 도전 대회에서 최신 기준 성능을 달성함으로써 기준 데이터셋에서 최고의 성능을 달성한다.

제안 방법

  • RGB 및 TIR 모달리티의 정규화된 분류 점수에 기반해 기여도를 동적으로 조정하는 결합 수준 융합 전략(DFAT)을 제안한다.
  • 분류 헤드에서 RGB 및 TIR 브랜치 간의 본질적 점수 격차를 보정하기 위해 동적 데 biases 메커니즘을 도입한다.
  • 분류 및 회귀 브랜치의 상대적 영향력을 균형 잡기 위해 스케일링 인자를 적용하여 융합의 강건성을 향상시킨다.
  • 10 프레임마다 대상 보간을 사용하여 선형 템플릿 업데이트 메커니즘을 구현하여 시간에 따라 정확한 대상 표현을 유지한다.
  • SiamRPN++ 백본을 RGB 데이터로만 훈련시키며, TIR 이미지를 세 개의 채널에 복제하여 추론 시 열화상 데이터 처리를 가능하게 한다.
  • 모달리티 신뢰도에 기반해 가중치를 적응적으로 학습하는 방식으로 단순 평균화를 피하는 융합 전략을 적용하여 분별력을 향상시킨다.
Figure 1: Qualitative comparison between our method (DFAT) and Baseline, with ground truth on the VOT-RGBT2019 [ 4 ] dataset. Respectively, frames sampled from video sequences bus6 , face1 , and woman89 are shown in the first, second and third rows of the figure. Here ’Baseline’ means SiamRPN++ (RGB
Figure 1: Qualitative comparison between our method (DFAT) and Baseline, with ground truth on the VOT-RGBT2019 [ 4 ] dataset. Respectively, frames sampled from video sequences bus6 , face1 , and woman89 are shown in the first, second and third rows of the figure. Here ’Baseline’ means SiamRPN++ (RGB

실험 결과

연구 질문

  • RQ1픽셀 수준, 특징 수준, 결합 수준 융합 전략 간의 정확도 및 강건성 측면에서 RGBT 추적에서의 성능 비교는 어떻게 되는가?
  • RQ2RGB 데이터로만 훈련된 트래커가 추론 시 결합 수준 융합을 통해 TIR를 융합할 경우, 더 뛰어난 성능을 달성할 수 있는가?
  • RQ3RGB 및 TIR 기여도에 대한 동적 가중치 적용이 외관 변화가 있는 상황에서 추적 정확도에 미치는 영향은 어떠한가?
  • RQ4장시간 시퀀스 동안 추적 안정성을 유지하는 데 선형 템플릿 업데이트 메커니즘이 얼마나 효과적인가?
  • RQ5복잡한 특징 수준 또는 픽셀 수준 융합 방법에 비해, 단순하면서도 적응 가능한 결합 수준 융합 전략이 RGBT 추적에서 더 뛰어난 성능을 낼 수 있는가?

주요 결과

  • 제안된 DFAT 트래커는 VOT-RGBT2020 데이터셋에서 최고의 EAO 점수 0.4178을 기록하여 새로운 SOTA를 수립하였다.
  • 편향 보정이 없는 DFAT의 변종(W/O bias)은 VOT-RGBT2020 도전 대회에서 우승하였으며, EAO 점수 0.4073를 기록하였고, 편향 보정 후 점수가 0.4178로 향상되었다.
  • VOT-RGBT2019 데이터셋에서 DFAT는 AUC 0.6652를 기록하여 SiamDW_T, JMMAC, mfDimp를 포함한 모든 비교 방법을 능가하였다.
  • 민감도 분석 결과, DFAT는 스케일링 인자에 대해 강건하며, 최적 성능은 0.47에서 달성되었고, 0.5에서 0.47로 조정했을 때 성능 향상은 0.04%에 불과하였다.
  • 제거 실험 결과, 동적 가중치와 템플릿 업데이트 모두 필수적임을 확인하였으며, 둘 중 하나를 제거할 경우 성능 저하가 심각하게 발생하였다.
  • GTOT, VOT-RGBT2019, VOT-RGBT2020에서의 광범위한 실험을 통해 DFAT가 다양한 추적 시나리오에서 강건성과 일반화 능력을 입증하였다.
Figure 2: Illustration of the proposed DFAT method. Here ’Cls’ and ’Reg’ represent classification and regression branches respectively. The left part describes the architecture of the feature extraction network (’Net’). The outputs from three convolutional layers are first projected into a common sp
Figure 2: Illustration of the proposed DFAT method. Here ’Cls’ and ’Reg’ represent classification and regression branches respectively. The left part describes the architecture of the feature extraction network (’Net’). The outputs from three convolutional layers are first projected into a common sp

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.