Skip to main content
QUICK REVIEW

[논문 리뷰] Learning Disentangled Representation Implicitly via Transformer for Occluded Person Re-Identification

Mengxi Jia, Xinhua Cheng|arXiv (Cornell University)|2021. 07. 06.
Video Surveillance and Tracking Methods인용 수 8
한 줄 요약

이 논문은 오염된 사람 재식별을 위한 트랜스포머 기반의 정렬 불필요 프레임워크인 DRL-Net을 제안한다. 이는 정의되지 않은 의미적 구성 요소에 대한 전역적 추론을 통해 오염 관련 노이즈에서 정체성 관련 특징을 암묵적으로 분리한다. 의미적 선호도 객체 쿼리와 분리 제약 조건이 있는 대비 특징 학습 모듈을 활용함으로써, DRL-Net은 Occluded-DukeMTMC에서 최신 기술 수준을 상당한 격차로 초월한다.

ABSTRACT

Person re-identification (re-ID) under various occlusions has been a long-standing challenge as person images with different types of occlusions often suffer from misalignment in image matching and ranking. Most existing methods tackle this challenge by aligning spatial features of body parts according to external semantic cues or feature similarities but this alignment approach is complicated and sensitive to noises. We design DRL-Net, a disentangled representation learning network that handles occluded re-ID without requiring strict person image alignment or any additional supervision. Leveraging transformer architectures, DRL-Net achieves alignment-free re-ID via global reasoning of local features of occluded person images. It measures image similarity by automatically disentangling the representation of undefined semantic components, e.g., human body parts or obstacles, under the guidance of semantic preference object queries in the transformer. In addition, we design a decorrelation constraint in the transformer decoder and impose it over object queries for better focus on different semantic components. To better eliminate interference from occlusions, we design a contrast feature learning technique (CFL) for better separation of occlusion features and discriminative ID features. Extensive experiments over occluded and holistic re-ID benchmarks (Occluded-DukeMTMC, Market1501 and DukeMTMC) show that the DRL-Net achieves superior re-ID performance consistently and outperforms the state-of-the-art by large margins for Occluded-DukeMTMC.

연구 동기 및 목표

  • 보이는 신체 부위가 일관되지 않거나 정렬이 잘못되어 성능이 떨어지는 심한 오염 조건에서 사람 재식별 문제를 해결하기 위해.
  • 자세 추정이나 부위 검출과 같은 복잡한 정렬 작업 또는 외부 감독을 필요로 하지 않기 위해.
  • 트랜스포머 아키텍처의 자기주의 메커니즘을 활용해 정체성 관련 특징과 오염 관련 노이즈를 암묵적으로 분리하기 위해.
  • 대비 학습과 분리 제약 조건을 통해 특징을 학습하여 오염 간섭에 대한 강건성을 향상시키기 위해.

제안 방법

  • DRL-Net은 CNN에서 추출한 특징에 대해 전역적 추론을 수행하기 위해 비전 트랜스포머 인코더-디코더 아키텍처를 사용하여 정렬 없이 표현을 학습한다.
  • 의미적 선호도 객체 쿼리는 트랜스포머가 정의되지 않은 의미적 구성 요소(예: 신체 부위 및 오염)를 별개의 특징 표현으로 분리하도록 이끈다.
  • 디코더에서 객체 쿼리에 분리 제약 조건을 적용하여 각 의미적 구성 요소에 대해 별개의 초점을 맞추고 쿼리 간 간섭을 줄인다.
  • 대비 특징 학습(CFL) 모듈은 전역 표현과 국소 특징을 대비시켜 정체성 관련 및 정체성 비관련 특징 간의 분리를 향상시킨다.
  • 훈련 중 다양한 오염 패턴을 시뮬레이션하기 위해 데이터 증강 전략을 통합한다.
  • 모델은 ID 분류와 대비 손실을 동시에 최적화하며, 하이퍼파ram터 λ와 α가 대비 및 분리 구성 요소를 균형 조절한다.
Figure 1: Illustration of the proposed DRL-Net in occluded person re-ID: Person images often suffer from various occlusions with different visible and invisible body parts which greatly complicate image alignment and similarity computation. The proposed DRL-Net is alignment-free which exploits seman
Figure 1: Illustration of the proposed DRL-Net in occluded person re-ID: Person images often suffer from various occlusions with different visible and invisible body parts which greatly complicate image alignment and similarity computation. The proposed DRL-Net is alignment-free which exploits seman

실험 결과

연구 질문

  • RQ1외부 부위 감독이나 정렬 없이 트랜스포머 기반 아키텍처가 오염 노이즈에서 정체성 관련 특징을 암묵적으로 분리할 수 있는가?
  • RQ2자기주의를 통한 전역적 추론이 오염된 사람 이미지의 정의되지 않은 의미적 구성 요소 간 상호관계를 모델링하는 데 얼마나 효과적인가?
  • RQ3대비 특징 학습이 오염 특징과 분류 가능한 ID 특징 간의 분리를 얼마나 향상시키는가?
  • RQ4하이퍼파ram터 λ(대비 손실 가중치)와 α(분리 페널티)가 모델의 강건성과 성능에 미치는 영향은 어떠한가?

주요 결과

  • DRL-Net은 Occluded-DukeMTMC 벤치마크에서 최신 기술 수준의 성능을 달성하여 이전 방법들보다 큰 격차로 앞서 있다.
  • Occluded-DukeMTMC에서 DRL-Net은 Rank-1 정확도 85.6%와 mAP 72.1%를 기록하여 이전 최고 성능를 상당히 초월했다.
  • 절단 실험 결과, 객체 쿼리 수(Nq)를 늘릴수록 Nq=9까지 성능 향상이 이루어지며, 이후에는 수익 감소와 함께 추론 비용 증가가 발생한다.
  • 모델은 하이퍼파ram터 변화에 강건하다: 최적의 성능은 λ=1.0과 α=1.0에서 달성되며, 다양한 값 범위에서 안정적인 성능 유지를 보였다.
  • 시각화 결과, 정체성 비관련 객체 쿼리가 어떤 감독 없이도 히트맵에서 오염 영역을 자동으로 국소화함으로써 효과적인 분리가 이루어졌음을 확인했다.
  • 정성적 결과에서는 DRL-Net이 오염된 쿼리 간 동일한 보행자를 정확히 순위 매기지만, 베이스라인 모델은 오염 민감도로 인해 많은 오진을 유발한다.
Figure 2: The framework of the proposed DRL-Net: DRL-Net consists of three components. The first component is occluded sample augmentation that synthesizes person images by inserting various obstacles. The second component is semantic representation extraction and disentanglement that disentangles t
Figure 2: The framework of the proposed DRL-Net: DRL-Net consists of three components. The first component is occluded sample augmentation that synthesizes person images by inserting various obstacles. The second component is semantic representation extraction and disentanglement that disentangles t

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.