Skip to main content
QUICK REVIEW

[논문 리뷰] Dynamic Mobile-Former: Strengthening Dynamic Convolution with Attention and Residual Connection in Kernel Space

Seokju Yun, Youngmin Ro|arXiv (Cornell University)|2023. 04. 13.
Advanced Neural Network Applications인용 수 4
한 줄 요약

이 논문은 동적 컨볼루션을 입력에 무관한 커널과 입력에 의존하는 커널을 잔차 연결을 통해 커널 공간에서 통합함으로써 향상시키는 경량 시각 모델인 동적 모바일-포머(DMF)를 제안한다. 또한 경량 어텐션을 사용해 커널 선택을 안내한다. DMF는 PVT-Tiny의 1/4에 불과한 FLOPs로 ImageNet-1K에서 79.4%의 top-1 정확도를 달성하며, 검출 및 세그멘테이션 작업에서 효율성과 정확도 면에서 최신 기술을 초월한다.

ABSTRACT

We introduce Dynamic Mobile-Former(DMF), maximizes the capabilities of dynamic convolution by harmonizing it with efficient operators.Our Dynamic MobileFormer effectively utilizes the advantages of Dynamic MobileNet (MobileNet equipped with dynamic convolution) using global information from light-weight attention.A Transformer in Dynamic Mobile-Former only requires a few randomly initialized tokens to calculate global features, making it computationally efficient.And a bridge between Dynamic MobileNet and Transformer allows for bidirectional integration of local and global features.We also simplify the optimization process of vanilla dynamic convolution by splitting the convolution kernel into an input-agnostic kernel and an input-dependent kernel.This allows for optimization in a wider kernel space, resulting in enhanced capacity.By integrating lightweight attention and enhanced dynamic convolution, our Dynamic Mobile-Former achieves not only high efficiency, but also strong performance.We benchmark the Dynamic Mobile-Former on a series of vision tasks, and showcase that it achieves impressive performance on image classification, COCO detection, and instanace segmentation.For example, our DMF hits the top-1 accuracy of 79.4% on ImageNet-1K, much higher than PVT-Tiny by 4.3% with only 1/4 FLOPs.Additionally,our proposed DMF-S model performed well on challenging vision datasets such as COCO, achieving a 39.0% mAP,which is 1% higher than that of the Mobile-Former 508M model, despite using 3 GFLOPs less computations.Code and models are available at https://github.com/ysj9909/DMF

연구 동기 및 목표

  • 커널 공간 학습의 최적화 불안정성과 제한된 표현 능력을 해결하기 위해 동적 컨볼루션의 재고찰.
  • 엄격한 FLOP 및 지연 시간 제약 조건 하에서 경량 시각 모델의 효율성과 성능 향상.
  • 입력에 무관한 커널과 입력에 의존하는 커널을 분리함으로써 동적 컨볼루션의 효과적인 확장성 확보.
  • 경량 어텐션에서 유도된 전역적 맥락을 커널 어텐션 모듈에 통합해 더 나은 필터 선택을 위한 개선.
  • 모바일 및 실시간 시각 응용 프로그램에서 정확도-FLOPs 간의 우수한 트레이드오프 달성.

제안 방법

  • 컨볼루션 커널을 입력에 무관한 성분과 입력에 의존하는 성분으로 분해하여 최적화를 분리하고 커널 공간을 확장.
  • 입력에 의존하는 커널을 입력에 무관한 커널에 잔차 연결을 통해 연결함으로써 학습 안정성 향상 및 최적화 개선.
  • 커널 공간 용량을 증가시키기 위해 활성화 함수로 시그모이드를 사용하여 최적화 난이도를 보완하는 잔차 학습을 통한 균형 조절.
  • 경량 전역 어텐션을 사용해 주목할 만한 전역 특징을 추출하고, 이를 커널 어텐션 모듈의 입력으로 활용해 더 나은 필터 선택을 위한 개선.
  • 동적 잔차 컨볼루션과 그룹 컨볼루션, 디프스와이즈 분리 컨볼루션의 조합을 통해 계산 효율성 유지.
  • 학습 안정성 향상과 수렴 개선을 위해 온도 안내와 철저한 초기화(예: 정적 커널의 0초기화)를 활용.
Figure 3 : A Dynamic Residual Convolution layer. Firstly, the global average pooled features of the input and the first global token are concatenated. These concatenated features are then passed through a kernel attention module to generate attention scores ( $\pi$ ) and obtain dynamic convolution w
Figure 3 : A Dynamic Residual Convolution layer. Firstly, the global average pooled features of the input and the first global token are concatenated. These concatenated features are then passed through a kernel attention module to generate attention scores ( $\pi$ ) and obtain dynamic convolution w

실험 결과

연구 질문

  • RQ1커널 공간 분해와 잔차 학습을 통해 동적 컨볼루션의 안정성과 확장성을 향상시킬 수 있는가?
  • RQ2경량 어텐션에서 유도된 전역 맥락을 사용할 경우 동적 컨볼루션의 성능 향상은 어떻게 이루어지는가?
  • RQ3학습 안정성에 영향을 주지 않으면서 더 큰 커널 공간을 효과적으로 활용할 수 있는가?
  • RQ4활성화 함수(시그모이드 대비 소프트맥스)가 동적 컨볼루션의 성능과 최적화에 미치는 영향은 무엇인가?
  • RQ5제안된 동적 잔차 컨볼루션은 표준 동적 컨볼루션 대비 FLOPs, 파라미터 수, 정확도 측면에서 어떻게 비교되는가?

주요 결과

  • DMF-S는 ImageNet-1K에서 79.4%의 top-1 정확도를 기록했으며, 이는 PVT-Tiny보다 4.3% 높고 FLOPs는 1/4 뿐이다.
  • DMF-S는 COCO 인스턴스 세그멘테이션에서 39.0%의 mAP를 달성했으며, Mobile-Former 508M 모델을 1% 뛰어넘었고, 계산량은 3 GFLOPs 적게 사용했다.
  • 정적 커널 8개와 시그모이드 기반 어텐션 점수를 사용한 모델은 커널 수가 적거나 소프트맥스 활성화를 사용한 모델보다 높은 73.6%의 top-1 정확도를 기록했다.
  • 정적 커널의 0초기화가 무작위 초기화보다 성능을 향상시켜 커널 공간에서의 잔차 학습의 이점이 확인되었다.
  • 제거 분석 결과, 시그모이드 활성화 함수와 잔차 연결이 모델 용량 향상과 학습 안정성 향상에 크게 기여하는 것으로 확인되었다.
  • DMF는 강력한 정확도-FLOPs 트레이드오프를 보이며, 실시간 및 모바일 시각 응용 분야에서 경량 CNN 및 비전 트랜스포머 변형보다 뛰어난 성능을 보였다.
Dynamic Mobile-Former: Strengthening Dynamic Convolution with Attention and Residual Connection in Kernel Space

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.