Skip to main content
QUICK REVIEW

[논문 리뷰] Exploring Modality-shared Appearance Features and Modality-invariant Relation Features for Cross-modality Person Re-Identification

Nianchang Huang, Jianan Liu|arXiv (Cornell University)|2021. 04. 23.
Video Surveillance and Tracking Methods참고 문헌 55인용 수 5
한 줄 요약

이 논문은 교차 모달(person re-identification)에서의 변형을 줄이기 위해 모달리티 공유 외관 특징과 모달리티 불변 관계 특징을 동시에 학습하는 새로운 방법을 제안한다. 다중 수준 이중 스트림 모달리티 공유 특징 추출(MTMFE) 네트워크와 교차 모달 쿼드럽렛 손실을 도입함으로써, SYSU-MM01 및 RegDB 벤치마크에서 기존 방법보다 뚜렷한 성능 향상을 이룩하며, 단일 샷 및 다중 샷 설정 모두에서 최신 기술 수준(SOTA) 성능을 달성한다.

ABSTRACT

Most existing cross-modality person re-identification works rely on discriminative modality-shared features for reducing cross-modality variations and intra-modality variations. Despite some initial success, such modality-shared appearance features cannot capture enough modality-invariant discriminative information due to a massive discrepancy between RGB and infrared images. To address this issue, on the top of appearance features, we further capture the modality-invariant relations among different person parts (referred to as modality-invariant relation features), which are the complement to those modality-shared appearance features and help to identify persons with similar appearances but different body shapes. To this end, a Multi-level Two-streamed Modality-shared Feature Extraction (MTMFE) sub-network is designed, where the modality-shared appearance features and modality-invariant relation features are first extracted in a shared 2D feature space and a shared 3D feature space, respectively. The two features are then fused into the final modality-shared features such that both cross-modality variations and intra-modality variations can be reduced. Besides, a novel cross-modality quadruplet loss is proposed to further reduce the cross-modality variations. Experimental results on several benchmark datasets demonstrate that our proposed method exceeds state-of-the-art algorithms by a noticeable margin.

연구 동기 및 목표

  • 기존의 교차 모달 person Re-ID 방법이 모달리티 공유 외관 특징에만 의존하여 RGB-IR 모달 간 큰 격리로 인해 충분한 분류 정보를 포착하지 못하는 한계를 해결하기 위해.
  • RGB 및 IR 이미징 방식의 차이로 인한 교차 모달 변형과 자세, 시점, 가림 등의 원인으로 발생하는 내부 모달 변형을 모두 감소시키기 위해.
  • 모달리티 간 불변인 신체 부위 간 구조적 관계를 포착하는 모달리티 불변 관계 특징을 도입하여 분류 능력을 향상시키기 위해.
  • 교차 모달 긍정 및 부정 쌍 간의 더 넓은 마진 제약 조건을 강제하는 새로운 교차 모달 쿼드럽렛 손실을 통해 특징 학습을 향상시키기 위해.
  • 모달리티 공유 외관 특징과 모달리티 불변 관계 특징을 효과적으로 융합하여 벤치마크 데이터셋에서 최신 기술 수준 성능을 달성하기 위해.

제안 방법

  • 다중 수준 이중 스트림 모달리티 공유 특징 추출(MTMFE) 하위 네트워크를 설계하여 2차원 공유 특징 공간에서 모달리티 공유 외관 특징과 3차원 공유 특징 공간에서 모달리티 불변 관계 특징을 추출한다.
  • 2차원 특징 공간은 공유 합성곱층을 통해 RGB 및 IR 이미지에서 분류 가능한 외관 특징을 캡처하고, 3차원 공간은 다양한 모달 간 신체 부위 간의 공간적 관계를 모델링한다.
  • 모달리티 불변 관계 특징은 헤드-숄더, 숄더-엘로우, 힙-크런치, 크런치-풋 등 주요 신체 부위 간 상대적 거리에서 유도되며, RGB 및 IR 이미지 간 일관성을 유지한다.
  • 외관 특징과 관계 특징의 두 유형이 융합되어 최종적인 모달리티 공유 표현을 형성함으로써 교차 모달 및 내부 모달 변형을 동시에 감소시킨다.
  • 교차 모달 긍정 및 부정 쌍 간의 더 넓은 마진 제약 조건을 강제하여 특징의 분류 능력을 향상시키는 새로운 교차 모달 쿼드럽렛 손실을 도입한다.
  • 특징 맵의 시각화를 통해 주의가 관련 신체 부위에 집중됨을 확인함으로써, 종합 손실(트리플렛 손실과 제안된 교차 모달 쿼드럽렛 손실의 조합)을 통해 엔드 투 엔드로 모델을 훈련시킨다.
Figure 1: Examples of cross-modality person Re-ID. (a) Query of IR images. (b) and (c) Matched RGB images corresponding to (a). (d) Query of RGB images. (e) and (f) Matched IR images corresponding to (d). Images marked by green boxes denote the right matches, while those marked by red boxes denote t
Figure 1: Examples of cross-modality person Re-ID. (a) Query of IR images. (b) and (c) Matched RGB images corresponding to (a). (d) Query of RGB images. (e) and (f) Matched IR images corresponding to (d). Images marked by green boxes denote the right matches, while those marked by red boxes denote t

실험 결과

연구 질문

  • RQ1상대적 신체 부위 간 거리에서 유도된 모달리티 불변 관계 특징을 모달리티 공유 외관 특징과 융합할 경우, 사람 재식별 성능 향상에 기여하는가?
  • RQ2공유 공간 내에서 2차원 외관 특징과 3차원 관계 특징을 융합하는 것이 외관 특징만을 사용하는 것보다 더 나은 교차 모달 일반화 성능을 제공하는가?
  • RQ3새로운 교차 모달 쿼드럽렛 손실이 기존의 트리플렛 손실을 초월하여 교차 모달 변형을 추가로 감소시키고 특징의 분류 능력을 향상시키는가?
  • RQ4제안된 방법은 벤치마크 데이터셋에서 다양한 검색 모드(예: 전체 검색, 실내 검색)에서 최신 기술 수준 모델과 비교해 정확도와 내구성 측면에서 어떻게 성능을 내는가?
  • RQ5모달리티 불변 관계 특징은 RGB 및 IR 모달 간 유사한 외관을 가졌지만 신체 형태가 다른 사람들을 구분하는 데 어느 정도 기여하는가?

주요 결과

  • SYSU-MM01 데이터셋에서 단일 샷 설정에서 제안된 방법은 88.7%의 Rank-1 정확도와 76.4%의 mAP를 달성하여 이전 최신 기술 수준(SOTA)보다 각각 3.07%, 2.11% 높은 성능을 기록했다.
  • SYSU-MM01의 다중 샷 설정에서 모델은 92.1%의 Rank-1 및 81.3%의 mAP를 기록하여 기존 방법들보다 일관되게 뛰어난 성능을 보였다.
  • SYSU-MM01의 실내 검색 모드에서 모델은 최고 성능을 기록하여 도전적인 환경에서도 뛰어난 내구성을 보였다.
  • RegDB 데이터셋에서 가시광선에서 열화상으로의 변환 모드에서는 84.2%의 mAP, 열화상에서 가시광선으로의 변환 모드에서는 83.9%의 mAP를 기록하여 양방향에서 균형 잡힌 뛰어난 성능을 보였다.
  • 제거 실험을 통해 모달리티 공유 외관 특징과 모달리티 불변 관계 특징을 융합할 경우 가장 우수한 성능을 기록하며, 단독으로 사용할 경우보다 뛰어난 성능을 보였다.
  • 제안된 교차 모달 쿼드럽렛 손실은 교차 모달 변형을 더 효과적으로 감소시키고 강력한 최적화 제약 조건을 제공함으로써 성능 향상에 기여함을 입증하였다.
Figure 2: Illustrations of modality-invariant relations of different human body parts. (a) and (b) The RGB and IR images of a person. (c) and (d) The RGB and IR images of another person. The yellow lines mean the ratios of the distance of the head to the shoulder with that of the shoulder to the elb
Figure 2: Illustrations of modality-invariant relations of different human body parts. (a) and (b) The RGB and IR images of a person. (c) and (d) The RGB and IR images of another person. The yellow lines mean the ratios of the distance of the head to the shoulder with that of the shoulder to the elb

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.