Skip to main content
QUICK REVIEW

[논문 리뷰] DaViT: Dual Attention Vision Transformers

Mingyu Ding, Bin Xiao|arXiv (Cornell University)|2022. 04. 07.
Advanced Neural Network Applications인용 수 11
한 줄 요약

DaViT는 시각 트랜스포머에 이중 주의 메커니즘을 도입하여 공간적이고 채널 기반의 자기주의를 동시에 모델링함으로써 지역적인 세밀한 표현과 전역적 맥락 표현을 효율적으로 포착한다. 공간 토큰과 전치된 채널 토큰에 대해 자기주의를 적용하고 토큰을 그룹화하여 순차 길이에 대해 선형 복잡도를 유지함으로써, DaViT는 계산 효율성을 유지하면서도 최신 기술 수준의 정확도를 달성한다—ImageNet-1K에서 84.6%의 top-1 정확도를 기록하며 파arameter 수는 87.9M이다.

ABSTRACT

In this work, we introduce Dual Attention Vision Transformers (DaViT), a simple yet effective vision transformer architecture that is able to capture global context while maintaining computational efficiency. We propose approaching the problem from an orthogonal angle: exploiting self-attention mechanisms with both "spatial tokens" and "channel tokens". With spatial tokens, the spatial dimension defines the token scope, and the channel dimension defines the token feature dimension. With channel tokens, we have the inverse: the channel dimension defines the token scope, and the spatial dimension defines the token feature dimension. We further group tokens along the sequence direction for both spatial and channel tokens to maintain the linear complexity of the entire model. We show that these two self-attentions complement each other: (i) since each channel token contains an abstract representation of the entire image, the channel attention naturally captures global interactions and representations by taking all spatial positions into account when computing attention scores between channels; (ii) the spatial attention refines the local representations by performing fine-grained interactions across spatial locations, which in turn helps the global information modeling in channel attention. Extensive experiments show our DaViT achieves state-of-the-art performance on four different tasks with efficient computations. Without extra data, DaViT-Tiny, DaViT-Small, and DaViT-Base achieve 82.8%, 84.2%, and 84.6% top-1 accuracy on ImageNet-1K with 28.3M, 49.7M, and 87.9M parameters, respectively. When we further scale up DaViT with 1.5B weakly supervised image and text pairs, DaViT-Gaint reaches 90.4% top-1 accuracy on ImageNet-1K. Code is available at https://github.com/dingmyu/davit.

연구 동기 및 목표

  • 시각 트랜스포머에서 전역 맥락 모델링와 계산 효율성 사이의 상충 관계를 해결하기 위해.
  • 공간 해상도에 대해 제곱 복잡도 없이 전역 장거리 의존성을 포착할 수 있는 메커니즘을 설계하기 위해.
  • 공간적 자기주의와 채널적 자기주의가 표현 학습을 향상시키는 데 기여하는 상호보완적 역할을 탐색하기 위해.
  • 이중 주의가 더 적은 파라미터 수와 FLOPs로 표준 국소 또는 전역 자기주의를 능가할 수 있음을 보여주기 위해.
  • 피드포워드 네트워크(FFNs)를 제거하고 더 깊은 주의 중심 아키텍처를 도입했을 때의 효과를 평가하기 위해.

제안 방법

  • 공간 토큰(패치 기반)과 채널 토큰(특징 기반)의 두 가지 유형의 토큰을 도입하여, 채널 차원을 따라 주의를 적용함으로써 전체 이미지의 전역 맥락을 모델링한다.
  • 전치된 특징 맵에 자기주의를 적용하여 채널 토큰을 생성하며, 각 채널 토큰은 모든 공간 위치를 통해 전체 이미지를 전역적으로 표현한다.
  • 순서 길이에 대해 선형 계산 복잡도를 유지하기 위해 공간 토큰과 채널 토큰을 순서 차원에 따라 그룹화한다.
  • 이중 주의 블록 내에서 국소 표현과 전역 표현을 개선하기 위해 공간 윈도우 자기주의와 채널 그룹 자기주의를 번갈아 적용한다.
  • FFNs를 제거하고 더 깊은 이중 주의 블록 스택으로 대체하여 파라미터 효율적이고 경량화된 설계를 구현함으로써 FLOP 예산을 충족시킨다.
  • 계산 비용을 줄이면서도 전역 모델링 능력을 유지하기 위해 그룹화된 쿼리와 키를 사용한 채널 그룹 자기주의를 적용한다.

실험 결과

연구 질문

  • RQ1공간적 및 채널적 자기주의를 동시에 모델링함으로써 계산 비용을 증가시키지 않고도 전역 맥락 모델링을 향상시킬 수 있는가?
  • RQ2정확도와 효율성 측면에서 이중 주의는 표준 윈도우 또는 전역 자기주의와 어떻게 비교되는가?
  • RQ3피드포워드 네트워크가 제거된 순수 주의 아키텍처가 이미지 분류에서 강력한 성능을 낼 수 있는가?
  • RQ4전역 특징 표현을 기반으로 작동하는 채널 자기주의는 기존의 SE나 ECA 모듈보다 성능이 뛰어나지 않는가?
  • RQ5이중 주의 메커니즘이 최신 기술 수준 성능에 도달하기 위해 필요한 레이어 수를 얼마나 줄일 수 있는가?

주요 결과

  • DaViT-Base는 파라미터 수 87.9M, FLOPs 15.2 GFLOPs로 ImageNet-1K에서 84.6%의 top-1 정확도를 기록하며 유사한 FLOP 수준의 이전 최신 기술 수준 메서드들을 능가한다.
  • DaViT-Tiny는 파라미터 수 28.3M, FLOPs 4.5 GFLOPs로 ImageNet-1K에서 82.8%의 top-1 정확도를 기록하며 뛰어난 효율-정확도 트레이드오프를 보여준다.
  • FFNs가 제거된 순수 주의 아키텍처인 DaViT의 변종은 ImageNet-1K에서 82.5%의 top-1 정확도를 기록하며, 윈도우 주의 및 채널 주의 기반 베이스라인보다 1.5–1.7% 높은 성능을 낸다.
  • 15억 개의 약한 감독을 받은 이미지-텍스트 쌍으로 사전학습된 DaViT-Giant는 ImageNet-1K에서 90.4%의 top-1 정확도를 달성하여 뛰어난 확장성을 보여준다.
  • DaViT 모델은 스위니 트랜스포머(Swin Transformer)보다 높은 처리량을 기록한다(예: V100에서 토이 모델 기준 1059 대 1024 samples/sec), 이는 더 나은 추론 효율성을 의미한다.
  • 채널 자기주의를 SE나 ECA 블록으로 대체할 경우 정확도가 1.6% 감소함을 확인하여, 제안된 주의 메커니즘이 전역 특징 융합에 더 효과적임을 입증한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.