Skip to main content
QUICK REVIEW

[논문 리뷰] Separable Self and Mixed Attention Transformers for Efficient Object Tracking

Goutam Yelluru Gopal, Maria A. Amer|arXiv (Cornell University)|2023. 09. 07.
Video Surveillance and Tracking Methods인용 수 5
한 줄 요약

이 논문은 경량화된 시각적 객체 추적기인 SMAT를 제안한다. 이는 백본에 분리 가능한 혼합 주의(attention)를, 헤드에 분리 가능한 자기 주의(self-attention)를 통합하여 효율적이고 고정확도의 추적을 가능하게 한다. 분리 가능한 혼합 주의를 통해 템플릿과 검색 영역의 특징을 융합하고, 효율적인 자기 주의를 통해 장거리 의존성을 모델링함으로써, SMAT는 CPU에서 37 fps, GPU에서 158 fps의 성능을 기록하며 파rameter 수가 3.8M에 불과하여 6개의 벤치마크에서 최신 경량 추적기들을 능가한다. 특히 GOT10k-test에서 AO 지표로 7.9% 향상된 성능을 기록한다.

ABSTRACT

The deployment of transformers for visual object tracking has shown state-of-the-art results on several benchmarks. However, the transformer-based models are under-utilized for Siamese lightweight tracking due to the computational complexity of their attention blocks. This paper proposes an efficient self and mixed attention transformer-based architecture for lightweight tracking. The proposed backbone utilizes the separable mixed attention transformers to fuse the template and search regions during feature extraction to generate superior feature encoding. Our prediction head performs global contextual modeling of the encoded features by leveraging efficient self-attention blocks for robust target state estimation. With these contributions, the proposed lightweight tracker deploys a transformer-based backbone and head module concurrently for the first time. Our ablation study testifies to the effectiveness of the proposed combination of backbone and head modules. Simulations show that our Separable Self and Mixed Attention-based Tracker, SMAT, surpasses the performance of related lightweight trackers on GOT10k, TrackingNet, LaSOT, NfS30, UAV123, and AVisT datasets, while running at 37 fps on CPU, 158 fps on GPU, and having 3.8M parameters. For example, it significantly surpasses the closely related trackers E.T.Track and MixFormerV2-S on GOT10k-test by a margin of 7.9% and 5.8%, respectively, in the AO metric. The tracker code and model is available at https://github.com/goutamyg/SMAT

연구 동기 및 목표

  • CPU와 같은 자원이 제한된 환경에서 표준 트랜스포머 기반 추적기의 계산 비용이 높은 문제를 해결하기 위해.
  • 백본 및 헤드 모듈 모두에 트랜스포머 기반 모델링을 적용할 수 있도록 경량 시아모이스 추적에 기여하기 위해.
  • 분리 가능한 혼합 주의를 통해 템플릿과 검색 영역 간의 특징 융합을 개선하여 더 나은 표현 학습을 가능하게 하기 위해.
  • 실시간 추론 속도(≥30 fps)를 유지하면서도 경량 추적에서 최고 성능을 달성하기 위해.
  • 분리 가능한 주의 메커니즘이 성능을 저하시키지 않고 표준 주의를 효과적으로 대체할 수 있음을 입증하기 위해.

제안 방법

  • 백본은 분리 가능한 혼합 주의 블록을 사용하는 하이브리드 CNN-ViT 아키텍처를 사용하여 템플릿 및 검색 영역의 특징을 동시에 추출하고 융합한다.
  • 분리 가능한 혼합 주의는 표준 행렬 곱셈을 요소별 연산으로 대체하여 교차 모odal 특징 융합 시 계산 비용을 감소시킨다.
  • 예측 헤드는 융합된 특징 내 장거리 의존성을 모델링하기 위해 분리 가능한 자기 주의 블록을 사용한다.
  • 이 아키텍처는 경량 추적기에서 처음으로 백본 및 헤드 모듈 모두에 트랜스포머 기반 구성 요소를 통합한다.
  • 모델은 표준 추적 손실 함수를 사용해 엔드 투 엔드로 훈련되며, CPU 배포를 최적화한 추론이 수행된다.
  • 제거 분석(ablation studies)을 통해 분리 가능한 주의 및 백본-헤드 조합 설계의 효과를 검증한다.
Figure 1 : Proposed SMAT architecture. The separable mixed attention Vision Transformer-based backbone jointly performs feature extraction and fusion of template and search regions. The separable transformer-based head models long-range dependencies within the fused features to predict accurate boun
Figure 1 : Proposed SMAT architecture. The separable mixed attention Vision Transformer-based backbone jointly performs feature extraction and fusion of template and search regions. The separable transformer-based head models long-range dependencies within the fused features to predict accurate boun

실험 결과

연구 질문

  • RQ1백본에서 표준 혼합 주의를 대체하기 위해 분리 가능한 혼합 주의가 계산 비용을 줄이면서도 특징 융합 품질을 유지하는 데 효과적인가?
  • RQ2헤드 모듈에서 분리 가능한 자기 주의가 경량 추적에서 컨볼루션 헤드와 비교해 유사하거나 더 높은 성능을 낼 수 있는가?
  • RQ3경량 추적기에서 트랜스포머 기반 백본과 헤드 모듈을 결합하면 실시간 속도를 유지하면서도 정확도가 향상되는가?
  • RQ4다양한 벤치마크에서 제안된 SMAT는 최신 경량 추적기들과 비교해 정확도 및 추론 속도 측면에서 어떤가?
  • RQ5백본에서의 특징 융합이 부분적 가림, 조명 변화, 변형과 같은 어려운 속성에 대한 강건성에 어떤 영향을 미치는가?

주요 결과

  • SMAT는 CPU에서 37 fps, GPU에서 158 fps를 기록하여 제한된 하드웨어 환경에서도 실시간 성능을 확보하며, 파라미터 수가 3.8M에 불과하다.
  • GOT10k-test에서 SMAT는 E.T.Track보다 AO 지표로 7.9% 향상되었고, MixFormerV2-S보다 5.8% 향상되었다.
  • LaSOT에서 SMAT는 14개 속성 중 8개에서 최고 성능, 5개에서 두 번째로 높은 성능을 기록했으며, 조명 변화 속성에서 MixFormerV2-S보다 AUC가 4.3% 높았다.
  • 부분적 가림(3.2% 높은 AUC)과 배경 혼잡성(2.7% 높은 AUC)에 대해 뚜렷한 성능 향상을 보였다.
  • 주의 시각화 결과, 부분적 가림 상황에서도 관련 시각적 단서에 주의를 집중함으로써 정확도가 향상됨을 확인했다.
  • 제거 분석 결과, 백본에서의 분리 가능한 혼합 주의와 헤드에서의 분리 가능한 자기 주의 조합이 뛰어난 정확도-속도 트레이드오프를 제공함을 입증했다.
Figure 2 : Proposed Separable Mixed Attention ViT block. The qkv-proj denotes the set of three $1\times 1$ convolutional filters to generate the Query , Key , and Value for attention computation. The mixed attention output is passed through a $1\times 1$ convolutional ffn-out block to generate the o
Figure 2 : Proposed Separable Mixed Attention ViT block. The qkv-proj denotes the set of three $1\times 1$ convolutional filters to generate the Query , Key , and Value for attention computation. The mixed attention output is passed through a $1\times 1$ convolutional ffn-out block to generate the o

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.