Skip to main content
QUICK REVIEW

[논문 리뷰] D$^3$ETR: Decoder Distillation for Detection Transformer

Xiaokang Chen, Jiahui Chen|arXiv (Cornell University)|2022. 11. 17.
Advanced Neural Network Applications인용 수 5
한 줄 요약

이 논문은 D³ETR를 제안하며, 순서가 없는 트랜스포머 디코더 출력 문제를 해결하기 위해 혼합 매칭 전략인 MixMatcher를 도입한 DETR 기반 객체 검출기용 지식 증류 방법이다. 교사 모델의 예측 및 어텐션 맵을 학생 모델로 증류함으로써 D³ETR는 최신 성능을 달성하였으며, 12 에포크 훈련에서 Conditional DETR-R50-C5의 성능을 7.8 mAP 향상시켰다.

ABSTRACT

While various knowledge distillation (KD) methods in CNN-based detectors show their effectiveness in improving small students, the baselines and recipes for DETR-based detectors are yet to be built. In this paper, we focus on the transformer decoder of DETR-based detectors and explore KD methods for them. The outputs of the transformer decoder lie in random order, which gives no direct correspondence between the predictions of the teacher and the student, thus posing a challenge for knowledge distillation. To this end, we propose MixMatcher to align the decoder outputs of DETR-based teachers and students, which mixes two teacher-student matching strategies, i.e., Adaptive Matching and Fixed Matching. Specifically, Adaptive Matching applies bipartite matching to adaptively match the outputs of the teacher and the student in each decoder layer, while Fixed Matching fixes the correspondence between the outputs of the teacher and the student with the same object queries, with the teacher's fixed object queries fed to the decoder of the student as an auxiliary group. Based on MixMatcher, we build extbf{D}ecoder extbf{D}istillation for extbf{DE}tection extbf{TR}ansformer (D$^3$ETR), which distills knowledge in decoder predictions and attention maps from the teachers to students. D$^3$ETR shows superior performance on various DETR-based detectors with different backbones. For example, D$^3$ETR improves Conditional DETR-R50-C5 by $ extbf{7.8}/ extbf{2.4}$ mAP under $12/50$ epochs training settings with Conditional DETR-R101-C5 as the teacher.

연구 동기 및 목표

  • 순서가 없는 디코더 출력으로 인해 표준화된 지식 증류(KD) 방법이 부족한 DETR 기반 검출기의 문제를 해결하기 위해.
  • 디코더 출력이 순열 불변성인 경우 교사 및 학생 예측 간의 매칭 문제를 극복하기 위해.
  • DETR 아키텍처의 트랜스포머 디코더에 특화된 강력하고 효과적인 KD 프레임워크를 개발하기 위해.
  • DETR 기반 객체 검출 모델의 지식 증류를 위한 새로운 베이스라인과 레시피를 수립하기 위해.

제안 방법

  • 각 디코더 레이어별로 이원적 매칭을 기반으로 하는 적응형 매칭과 공유된 객체 쿼리를 보조 입력으로 사용하는 고정 매칭을 조합한 하이브리드 매칭 전략인 MixMatcher를 제안한다.
  • 최적의 이원적 할당 기반으로 각 디코더 레이어에서 교사 및 학생 예측 간의 동적 정렬을 수행함으로써 적응형 매칭을 사용한다.
  • 교사의 고정된 객체 쿼리를 학생의 디코더에 보조 그룹으로 입력하여 정합성을 높임으로써 고정 매칭을 구현한다.
  • 최종 예측 외에도 디코더 레이어 내의 자기 어텐션 및 교차 어텐션 맵에 대해서도 지식 증류를 수행한다.
  • 동일한 쿼리에서 유사한 교사 및 학생 출력 간 일관성 있는 감독을 확보하기 위해 고정 매칭에 제약 조건을 적용한다.
  • 학습 성능 향상을 위해 학생 모델를 교사 모델의 가중치로 초기화하는 유산 전략을 통합한다.
Figure 2 : Architecture of the proposed method. We propose the mixed teacher-student matching strategy that composes of two components, adaptative matching and fixed matching. We adopt two groups, where the first group feeds student queries to the decoder and the second feeds teacher queries to the
Figure 2 : Architecture of the proposed method. We propose the mixed teacher-student matching strategy that composes of two components, adaptative matching and fixed matching. We adopt two groups, where the first group feeds student queries to the decoder and the second feeds teacher queries to the

실험 결과

연구 질문

  • RQ1순서가 없는 디코더 출력 특성으로 인해 DETR 기반 검출기의 지식 증류를 효과적으로 적용할 수 있는 방법은 무엇인가?
  • RQ2DETR의 디코더에서 교사 및 학생 모델 간 지식 전달을 안정화하고 향상시키기 위해 가장 적합한 매칭 전략은 무엇인가?
  • RQ3예측 외에도 어텐션 맵에 대한 증류가 학생 모델의 성능 향상에 상당한 기여를 할 수 있는가?
  • RQ4적응형 매칭과 고정 매칭 전략을 조합할 경우 증류의 안정성과 정확도에 어떤 영향을 미치는가?
  • RQ5고정 매칭 메커니즘에서 보조 객체 쿼리와 제약 감독을 사용할 경우 어떤 영향을 미치는가?

주요 결과

  • Conditional DETR-R101-C5를 교사로 사용할 경우, 12 에포크 훈련에서 Conditional DETR-R50-C5의 성능을 7.8 mAP 향상시키며, 50 에포크 훈련에서는 2.4 mAP 향상된다.
  • MixMatcher에서 적응형 매칭과 고정 매칭의 조합이 각각 단독으로 사용할 경우보다 38.1 mAP를 기록하며 우월한 성능을 낸다.
  • 고정 매칭에 제약 조건을 추가하면 성능이 38.1에서 38.1 mAP로 향상되며, 제약 조건을 제거할 경우 성능이 37.0 mAP로 떨어진다.
  • 시각화 결과는 D³ETR로 훈련된 학생 모델가 교사와 유사한 정확한 공간 어텐션 패턴을 학습하는 것으로 확인되었다.
  • D³ETR는 COCO에서 40.2 mAP를 달성하였으며, 이는 전체 50 에포크 훈련을 수행한 Conditional DETR-R50-C5(40.9 mAP)와 유사한 성능이다.
  • D³ETR를 MGD와 조합하면 성능이 38.8 mAP로 향상되어 다른 KD 방법과의 호환성을 입증한다.
Figure 3 : Analysis of different components in Conditional DETR. We adopt ResNet- $101$ /ResNet- $50$ as the backbone and train the model for $12$ epochs (1 $\times$ schedule). Best viewed in color.
Figure 3 : Analysis of different components in Conditional DETR. We adopt ResNet- $101$ /ResNet- $50$ as the backbone and train the model for $12$ epochs (1 $\times$ schedule). Best viewed in color.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.