Skip to main content
QUICK REVIEW

[논문 리뷰] RelationTrack: Relation-aware Multiple Object Tracking with Decoupled Representation

En Yu, Zhuoling Li|arXiv (Cornell University)|2021. 05. 10.
Video Surveillance and Tracking Methods인용 수 8
한 줄 요약

RelationTrack는 최적화 모순을 해결하기 위해 검출 및 재식별(ReID) 표현을 전역적 맥락 해리화(GCD) 모듈을 통해 분리하는 새로운 온라인 다중 객체 추적 프레임워크를 제안한다. 동시에 타원형 어텐션을 활용한 가이드드 트랜스포머 인코더(GTE)를 사용하여 효율적으로 전역적 의미 관계를 모델링한다. 이는 MOT20에서 70.5%의 IDF1과 67.2%의 MOTA를 달성하여 최신 기준 성능(SOTA)을 기록한다.

ABSTRACT

Existing online multiple object tracking (MOT) algorithms often consist of two subtasks, detection and re-identification (ReID). In order to enhance the inference speed and reduce the complexity, current methods commonly integrate these double subtasks into a unified framework. Nevertheless, detection and ReID demand diverse features. This issue would result in an optimization contradiction during the training procedure. With the target of alleviating this contradiction, we devise a module named Global Context Disentangling (GCD) that decouples the learned representation into detection-specific and ReID-specific embeddings. As such, this module provides an implicit manner to balance the different requirements of these two subtasks. Moreover, we observe that preceding MOT methods typically leverage local information to associate the detected targets and neglect to consider the global semantic relation. To resolve this restriction, we develop a module, referred to as Guided Transformer Encoder (GTE), by combining the powerful reasoning ability of Transformer encoder and deformable attention. Unlike previous works, GTE avoids analyzing all the pixels and only attends to capture the relation between query nodes and a few self-adaptively selected key samples. Therefore, it is computationally efficient. Extensive experiments have been conducted on the MOT16, MOT17 and MOT20 benchmarks to demonstrate the superiority of the proposed MOT framework, namely RelationTrack. The experimental results indicate that RelationTrack has surpassed preceding methods significantly and established a new state-of-the-art performance, e.g., IDF1 of 70.5% and MOTA of 67.2% on MOT20.

연구 동기 및 목표

  • 통합 추적 프레임워크에서 검출과 ReID 간의 최적화 모순 문제를 해결하기 위해.
  • 국소적 특징을 초월한 전역적 의미 관계 모델링을 통해 추적 정확도를 향상시키기 위해.
  • 장거리 의존성 모델링을 유지하면서도 전역 어텐션 기반 메커니즘의 계산 비용을 줄이기 위해.
  • 온라인 MOT에서 효율적인 관계 추론을 위한 경량이면서도 강력한 모듈을 개발하기 위해.

제안 방법

  • 최적화 갈등을 완화하기 위해 공유 특징을 검출 전용 및 ReID 전용 임베딩으로 분리하는 자기 동기화 모듈인 전역적 맥락 해리화(GCD)를 도입한다.
  • 변형 가능 어텐션과 트랜스포머 인코더 아키텍처를 결합한 가이드드 트랜스포머 인코더(GTE)를 제안하여 계산량을 줄이고 전역적 구조적 관계를 효과적으로 포착한다.
  • 각 쿼리 노드에서 자가 적응형 키 샘플을 몇 개만 선택함으로써 복잡도를 O(n²)에서 O(n)으로 감소시킨다.
  • 검출과 ReID가 분리된 표현을 통해 독립적으로 최적화되는 이중 브랜치 특징 학습 전략을 활용한다.
  • 단일 네트워크 내에서 검출, ReID, 관계 추론의 엔드 투 엔드 학습을 가능하게 하는 공동 학습 프레임워크를 설계한다.
  • 성능과 계산 비용의 균형을 맞추기 위해 GTE 내 키 샘플 수를 최적화하며, 쿼리당 9개의 샘플을 최적값으로 선택한다.

실험 결과

연구 질문

  • RQ1검출과 ReID 표현을 분리함으로써 최적화 갈등을 줄이고 추적 정확도를 향상시킬 수 있는가?
  • RQ2경량 전역 관계 모델링 메커니즘이 다중 객체 추적에서 국소적 또는 전역 어텐션보다 우월한 성능을 낼 수 있는가?
  • RQ3변형 가능 어텐션 내 키 샘플 수는 추적 성능와 계산 효율성에 어떤 영향을 미치는가?
  • RQ4관계 인식 프레임워크는 가림이나 소형 타깃과 같은 도전적인 조건에서도 강건성을 유지할 수 있는가?

주요 결과

  • RelationTrack는 MOT20 벤치마크에서 새로운 최고 기록인 70.5%의 IDF1과 67.2%의 MOTA를 달성한다.
  • MOT17에서 GCD 모듈을 사용함으로써 IDF1이 73.3%에서 74.9%로 향상되어 최적화 갈등 완화 효과가 효과적으로 입증된다.
  • 9개의 키 샘플을 사용한 GTE 모듈이 최적의 균형을 이루며, MOT17에서 75.3%의 IDF1과 70.1%의 MOTA를 기록했고, 9개를 초과해도 성능 저하가 거의 없었다.
  • MOT16에서는 FairMOT 대비 3.0% 높은 IDF1, MOT17에서는 2.4% 높은 IDF1을 기록하여 최신 기준 성능을 확인한다.
  • 강건성 분석 결과, FairMOT와 CSTrack가 실패하는 부분 가림 상황에서도 RelationTrack는 성공적으로 타깃을 추적함을 확인하여 강력한 특징 표현 능력을 입증한다.
  • 시각화 결과 GCD가 검출 전용 특징에서 검출에 관련된 영역을 효과적으로 강조하고, ReID 전용 특징에서 신원에 관련된 부분을 잘 드러냄을 확인했다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.