[논문 리뷰] Learning Spatial-Temporal Regularized Correlation Filters for Visual Tracking
STRCF는 SRDCF에 시간적 규제를 도입하고 ADMM으로 해결하여 여러 벤치마크에서 SRDCF 대비 실시간 추적 속도와 향상된 정확도를 달성한다.
Discriminative Correlation Filters (DCF) are efficient in visual tracking but suffer from unwanted boundary effects. Spatially Regularized DCF (SRDCF) has been suggested to resolve this issue by enforcing spatial penalty on DCF coefficients, which, inevitably, improves the tracking performance at the price of increasing complexity. To tackle online updating, SRDCF formulates its model on multiple training images, further adding difficulties in improving efficiency. In this work, by introducing temporal regularization to SRDCF with single sample, we present our spatial-temporal regularized correlation filters (STRCF). Motivated by online Passive-Agressive (PA) algorithm, we introduce the temporal regularization to SRDCF with single sample, thus resulting in our spatial-temporal regularized correlation filters (STRCF). The STRCF formulation can not only serve as a reasonable approximation to SRDCF with multiple training samples, but also provide a more robust appearance model than SRDCF in the case of large appearance variations. Besides, it can be efficiently solved via the alternating direction method of multipliers (ADMM). By incorporating both temporal and spatial regularization, our STRCF can handle boundary effects without much loss in efficiency and achieve superior performance over SRDCF in terms of accuracy and speed. Experiments are conducted on three benchmark datasets: OTB-2015, Temple-Color, and VOT-2016. Compared with SRDCF, STRCF with hand-crafted features provides a 5 times speedup and achieves a gain of 5.4% and 3.6% AUC score on OTB-2015 and Temple-Color, respectively. Moreover, STRCF combined with CNN features also performs favorably against state-of-the-art CNN-based trackers and achieves an AUC score of 68.3% on OTB-2015.
연구 동기 및 목표
- Discriminative correlation filters (DCFs) for visual tracking의 경계 효과를 해결한다.
- 단일 프레임에서 업데이트되는 시공-시간 규제 STRCF를 제안한다.
- 닫힌 형식의 부분 문제를 갖는 효율적인 ADMM 기반 해를 개발한다.
- 큰 appearance variation에서도 실시간 속도를 유지하며 STRCF가 견고한 appearance 모델을 제공한다.
제안 방법
- SRDCF에 mu/2 * ||f - f_{t-1}||^2의 시간 규제 항을 도입하여 STRCF를 형성한다 (Eq. 2).
- Auxiliary variable g와 교대 업데이트를 통해 볼록한 STRCF 목적함수를 ADMM으로 해결한다.
- f-subproblem에서는 Parseval’s theorem과 Sherman–Morrison formula를 사용하여 Fourier domain에서 위치별로 해결하여 효율성을 확보한다 (Eq. 9–12).
- g-subproblem에서는 대각 구조를 활용한 닫힌 형식의 해를 얻는다 (Eq. 13).
- ADMM 페널티 매개변수 gamma를 반복적으로 업데이트한다 (Eq. 14).
- 프레임당 계산 복잡도는 O(DMN log(MN))이며 전체 비용은 O(DMN log(MN) NI)이다.
- 수렴 보장을 확립한다(볼록 문제; Eckstein–Bertsekas 조건)하고 경험적으로 두 번의 반복 수렴을 보인다.
실험 결과
연구 질문
- RQ1STRCF가 여러 학습 이미지에서 학습된 SRDCF 모델을 근사하면서도 더 높은 효율성을 유지할 수 있는가?
- RQ2시간 규제를 도입하는 것이 SRDCF와 비교했을 때 appearance variations 및 occlusions에 대한 강건성을 개선하는가?
- RQ3시간 규제 매개변수 mu가 추적 성능에 미치는 영향은 무엇인가?
- RQ4STRCF가 손으로 설계된 특징과 딥 특징을 사용하면서도 실시간 성능을 달성하고 경쟁력 있는 정확도를 유지할 수 있는가?
주요 결과
- STRCF는 OTB-2015 및 Temple-Color에서 SRDCF 대비 평균 OP 약 5.7%의 이득을 달성한다.
- STRCF는 핸드크래프트 특징으로 실시간으로 실행되며 약 30 FPS를 달성하고 STRCF(HOG)로는 31.5 FPS, STRCF(HOGCN)로는 24.3 FPS를 달성한다.
- 시간 규제 도입 STRCF는 강건한 업데이트를 제공하며 OV 및 OCC 속성에서 SRDCF 변형 대비 각각 최대 14.5% 및 5.7%의 이득을 제공한다.
- DeepSTRCF( CNN 특징을 갖춘 STRCF)는 OTB-2015에서 평균 OP 84.2%를 달성하여 DeepSRDCF 대비 7.4% 향상된다.
- VOT-2016에서 STRCF의 EAO는 0.279(STRCF) 및 0.313(DeepSTRCF)이며, CNN 보강 변형 중 DeepSTRCF가 더 높은 EAO를 보인다.
- Temple-Color에서 STRCF는 ECO-HC와 경쟁력이 있으며 DeepSTRCF가 해당 데이터셋에서 보고된 결과 중 최고 성능을 보인다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.