Skip to main content
QUICK REVIEW

[논문 리뷰] DETR for Crowd Pedestrian Detection

Matthieu Lin, Chuming Li|arXiv (Cornell University)|2020. 12. 12.
Advanced Neural Network Applications참고 문헌 39인용 수 38
한 줄 요약

PED-DETR은 군중에서 DETR이 성능이 떨어지는 이유를 분석하고 DQRF, V-Match, Fast-KM을 도입하여 CrowdHuman과 CityPersons에서 최첨단 결과를 달성한다.

ABSTRACT

Pedestrian detection in crowd scenes poses a challenging problem due to the heuristic defined mapping from anchors to pedestrians and the conflict between NMS and highly overlapped pedestrians. The recently proposed end-to-end detectors(ED), DETR and deformable DETR, replace hand designed components such as NMS and anchors using the transformer architecture, which gets rid of duplicate predictions by computing all pairwise interactions between queries. Inspired by these works, we explore their performance on crowd pedestrian detection. Surprisingly, compared to Faster-RCNN with FPN, the results are opposite to those obtained on COCO. Furthermore, the bipartite match of ED harms the training efficiency due to the large ground truth number in crowd scenes. In this work, we identify the underlying motives driving ED's poor performance and propose a new decoder to address them. Moreover, we design a mechanism to leverage the less occluded visible parts of pedestrian specifically for ED, and achieve further improvements. A faster bipartite match algorithm is also introduced to make ED training on crowd dataset more practical. The proposed detector PED(Pedestrian End-to-end Detector) outperforms both previous EDs and the baseline Faster-RCNN on CityPersons and CrowdHuman. It also achieves comparable performance with state-of-the-art pedestrian detection methods. Code will be released soon.

연구 동기 및 목표

  • DETR와 deformable DETR이 FPN을 가진 Faster-RCNN에 비해 군중 보행자 검출에서 왜 성능이 떨어지는지 평가한다.
  • DETR의 보행자 성능 향상을 위한 Dense Query and Rectified Attention Field decoder (DQRF)을 제안한다.
  • V-Match를 이용해 가시 영역 주석을 활용하여 DETR 기반 보행자 검출을 향상시킨다.
  • 혼잡한 데이터셋에서 DETR 학습을 가능하게 하기 위해 더 빠른 이분 매칭 (Fast-KM)을 도입한다.
  • CrowdHuman 및 CityPersons 벤치마크에서 최첨단 또는 경쟁력 있는 성능을 입증한다.

제안 방법

  • 혼잡한 보행자 장면에서 DETR의 디코더 동작을 분석하여 실패 모드를 파악한다.
  • Dense Query와 보정된 더 넓은 주의 필드를 가능하게 하는 DQRF 디코더를 개발한다.
  • RF (Rectified Attention Field)를 도입하여 디코더 계층 간의 교차 주의를 안정화한다.
  • 추가 비용 없이 계층 간에 전체 가시 보행자 영역과 가시 영역을 감독하기 위해 V-Match를 제안한다.
  • 학습 중 Hungarian 매칭 단계를 가속화하기 위해 Fast-KM을 구현한다.

실험 결과

연구 질문

  • RQ1원래의 DETR과 deformable DETR이 Faster-RCNN과 비교해 군중 보행자 검출에서 왜 성능이 떨어지는가?
  • RQ2Dense Query and Rectified Attention Field 디코더가 DETR의 군중 보행자 검출 향상에 기여할 수 있는가?
  • RQ3V-Match를 통한 가시 영역 주석 활용이 추가 비용 없이 엔드-투-엔드 보행자 검출을 개선하는가?
  • RQ4정확도 손실 없이 Fast-KM으로 이분 매칭 학습에서 얼마나 속도 향상을 달성할 수 있는가?

주요 결과

  • PED는 CrowdHuman 및 CityPersons에서 deformable DETR 및 Faster-RCNN 벤치마크 대비 개선을 달성한다.
  • 제안된 DQRF 디코더는 혼잡한 장면에서 DETR 기반 탐지기와 Faster-RCNN 사이의 격차를 크게 줄인다.
  • V-Match는 추가 비용 없이 가시 영역 감독 이득을 제공한다.
  • Fast-KM은 학습 중 이분 매칭에 최대 10배의 속도 향상을 제공합니다.
  • PED는 도전적인 벤치마크에서 최첨단 방법과 경쟁력 있는 결과를 달성한다.
  • Dense queries와 rectified attention fields는 혼잡하게 가려진 환경에서 누락 검출을 줄인다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.