Skip to main content
QUICK REVIEW

[논문 리뷰] AdaMixer: A Fast-Converging Query-Based Object Detector

Ziteng Gao, Limin Wang|arXiv (Cornell University)|2022. 03. 30.
Advanced Neural Network Applications인용 수 5
한 줄 요약

AdaMixer는 적응적인 3차원 특징 샘플링과 동적 MLP-Mixer 디코딩을 통해 쿼리 적응성을 향상시켜, 추가적인 주의 메커니즘 인코더나 명시적인 피라미드 네트워크 없이도 ResNeXt-Swin-S와 12-에포크 훈련으로 MS COCO에서 최신 기준 AP 51.3을 달성하는 빠르게 수렴하는 쿼리 기반 객체 검출기이다.

ABSTRACT

Traditional object detectors employ the dense paradigm of scanning over locations and scales in an image. The recent query-based object detectors break this convention by decoding image features with a set of learnable queries. However, this paradigm still suffers from slow convergence, limited performance, and design complexity of extra networks between backbone and decoder. In this paper, we find that the key to these issues is the adaptability of decoders for casting queries to varying objects. Accordingly, we propose a fast-converging query-based detector, named AdaMixer, by improving the adaptability of query-based decoding processes in two aspects. First, each query adaptively samples features over space and scales based on estimated offsets, which allows AdaMixer to efficiently attend to the coherent regions of objects. Then, we dynamically decode these sampled features with an adaptive MLP-Mixer under the guidance of each query. Thanks to these two critical designs, AdaMixer enjoys architectural simplicity without requiring dense attentional encoders or explicit pyramid networks. On the challenging MS COCO benchmark, AdaMixer with ResNet-50 as the backbone, with 12 training epochs, reaches up to 45.0 AP on the validation set along with 27.9 APs in detecting small objects. With the longer training scheme, AdaMixer with ResNeXt-101-DCN and Swin-S reaches 49.5 and 51.3 AP. Our work sheds light on a simple, accurate, and fast converging architecture for query-based object detectors. The code is made available at https://github.com/MCG-NJU/AdaMixer

연구 동기 및 목표

  • 다양한 객체에 대한 디코더 적응성 향상을 통해 쿼리 기반 객체 검출기의 느린 수렴과 제한된 성능 문제를 해결한다.
  • 추가적인 주의 메커니즘 인코더나 명시적인 피라미드 네크를 제거하여 아키텍처 복잡성을 줄인다.
  • 적응적인 공간적 및 스케일 기반 특징 샘플링과 동적 콘텐츠 혼합을 통해 쿼리 기반 디코딩을 향상시킨다.
  • 최소한의 훈련 시간과 단순한 데이터 증강 기법으로 높은 정확도를 달성하여 향후 연구의 강력한 베이스라인을 제공한다.

제안 방법

  • 각 쿼리가 추정된 오프셋을 사용해 공간 위치와 스케일 수준에서 특징을 동적으로 샘플링하는 적응적인 3차원 특징 샘플링을 도입한다.
  • 다중 스케일 특징 맵을 3차원 텐서 (H×W×C)로 표현하여 샘플링 중 공간적 및 스케일 수준의 동시 주의를 가능하게 한다.
  • 샘플된 특징에 대해 MLP-Mixer를 통해 적응적인 채널 및 공간 혼합을 수행하는 쿼리 가로운 동적 커널을 적용한다.
  • 공간-스케일 샘플링과 이후 특징 혼합 과정을 안내하기 위해 학습 가능한 쿼리 임베딩을 사용한다.
  • FPN나 Transformer 인코더와 같은 추가 모듈이 필요 없이 엔드 투 엔드로 훈련 가능한 디코더를 설계한다.
  • 다양한 기반 특징의 효율적이고 기울기 기반의 특징 보간을 위해 PyTorch의 grid_sample를 사용해 적응적 샘플링을 구현한다.
Figure 1 : Convergence curves of our AdaMixer, DETR, Deformable DETR and Sparse R-CNN with ResNet-50 as the backbone on MS COCO minival set.
Figure 1 : Convergence curves of our AdaMixer, DETR, Deformable DETR and Sparse R-CNN with ResNet-50 as the backbone on MS COCO minival set.

실험 결과

연구 질문

  • RQ1복잡한 보조 네트워크에 의존하지 않고도 쿼리 기반 객체 검출기가 더 빠른 수렴과 높은 정확도를 달성할 수 있는가?
  • RQ2적응적인 3차원 특징 샘플링은 다양한 객체 크기와 위치에서 쿼리 표현을 어떻게 향상시키는가?
  • RQ3동적 MLP-Mixer 디코딩은 아키텍처의 단순성 유지와 함께 특징 혼합을 얼마나 향상시킬 수 있는가?
  • RQ4단지 12 에포크의 훈련과 최소한의 증강 기법으로도 기존 최신 기준 모델을 초월할 수 있는가?
  • RQ5RoIAlign과 가변성 컨볼루션에 비해 다중 스케일 모델링과 특징 샘플링의 유연성 측면에서 제안된 방법은 어떻게 비교되는가?

주요 결과

  • ResNet-50를 사용한 AdaMixer는 단지 12개의 훈련 에포크와 무작위 플립 증강만으로 MS COCO minival에서 45.0 AP를 기록하며, DETR 및 Deformable DETR를 능가한다.
  • 3배 훈련 스케줄과 더 강력한 증강 기법을 사용한 AdaMixer는 Swin-S 백본을 사용해 COCO test-dev에서 51.3 AP를 달성하여 이전의 쿼리 기반 검출기들을 뛰어넘었다.
  • ResNet-50 기반으로 작은 객체(APs)에 대해 27.9 AP를 기록하여 미세한 감지 성능이 뛰어나다는 것을 확인했다.
  • ResNeXt-101-DCN 기반 AdaMixer는 49.5 AP를 기록하여 더 큰 백본에서의 뛰어난 확장성과 성능을 입증했다.
  • 더 높은 파라미터 수에도 불구하고, DETR 및 Sparse R-CNN보다 더 빠른 수렴과 더 나은 추론 속도(15 FPS, V100 기준)를 보였다.
  • 시각화 결과, 쿼리와 단계에 따라 샘플링 포인트의 분포가 다르며, 비단조화적인 확산 패턴을 보여 추론 정밀도 향상 가능성을 시사한다.
Figure 2 : 3D feature sampling process. A query first obtains sampling points in the 3D feature space and then perform 3D interpolation on these sampling points.
Figure 2 : 3D feature sampling process. A query first obtains sampling points in the 3D feature space and then perform 3D interpolation on these sampling points.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.