Skip to main content
QUICK REVIEW

[논문 리뷰] FLatten Transformer: Vision Transformer using Focused Linear Attention

Dongchen Han, Xuran Pan|arXiv (Cornell University)|2023. 08. 01.
CCD and CMOS Imaging Sensors인용 수 15
한 줄 요약

Focused Linear Attention를 도입하여 Softmax를 교체한 비전 트랜스포머에서 선형 복잡도와 함께 표현력을 향상시키기 위해 집중 매핑과 깊이wise 컨볼루션으로 특성 다양성을 회복합니다. 분류, 분할, 탐지 벤치마크에서 일관된 성능 향상을 실험적으로 입증합니다.

ABSTRACT

The quadratic computation complexity of self-attention has been a persistent challenge when applying Transformer models to vision tasks. Linear attention, on the other hand, offers a much more efficient alternative with its linear complexity by approximating the Softmax operation through carefully designed mapping functions. However, current linear attention approaches either suffer from significant performance degradation or introduce additional computation overhead from the mapping functions. In this paper, we propose a novel Focused Linear Attention module to achieve both high efficiency and expressiveness. Specifically, we first analyze the factors contributing to the performance degradation of linear attention from two perspectives: the focus ability and feature diversity. To overcome these limitations, we introduce a simple yet effective mapping function and an efficient rank restoration module to enhance the expressiveness of self-attention while maintaining low computation complexity. Extensive experiments show that our linear attention module is applicable to a variety of advanced vision Transformers, and achieves consistently improved performances on multiple benchmarks. Code is available at https://github.com/LeapLabTHU/FLatten-Transformer.

연구 동기 및 목표

  • 비전 Transformers에서 자기 주의의 높은 계산 비용을 해결한다.
  • 선형 주의와 Softmax 주의 간의 성능 격차를 줄인다.
  • 집중과 특징 다양성을 개선하는 메커니즘으로 선형 주의를 강화한다.
  • 여러 Vision Transformer 아키텍처에 적용 가능한 플러그인 모듈을 제공한다.

제안 방법

  • 간단한 포커스 매핑과 depthwise convolution(DWC)에 의한 랭크 복원을 결합한 Focused Linear Attention 모듈을 제안한다.
  • 쿼리/ 키 방향을 조정하여 주의 분포를 샤프하게 만드는 매핑 함수 fp로 Softmax를 근사한다.
  • 특성의 랭크를 회복하고 다양성을 높이기 위해 V에 추가적인 DWC를 적용한다.
  • 주의를 O = Sim(Q,K)V = fp(Q) fp(K)^T V + DWC(V)로 형식화한다.
  • 계산을 Q(K^T V)로 재배열하여 선형 시간 복잡도를 입증한다(= (QK^T)V 대신).
  • ImageNet, ADE20K, COCO에서 DeiT, PVT, PVT-v2, Swin, CSWin 백본에 대한 플러그인 모듈로 평가한다.

실험 결과

연구 질문

  • RQ1집중된 선형 주의가 비전 트랜스포머에서 선형 계산 비용으로 Softmax 주의에 필적하거나 더 높은 정확도를 달성할 수 있는가?
  • RQ2간단한 매핑 기반 포커스 조정과 depthwise convolution 기반 랭크 회복이 선형 주의의 표현력과 특징 다양성을 향상시키는가?
  • RQ3Focused Linear Attention 모듈이 주요 비전 트랜스포머 아키텍처 전반에 걸쳐 플러그인으로 넓게 호환되는가?
  • RQ4baseline 주의 대신 FLatten 주의를 적용했을 때 ImageNet-1K, ADE20K, COCO에서의 실증적 이득은 무엇인가?

주요 결과

  • Focused Linear Attention은 기본 선형 주의보다 향상되며 여러 모델에서 Softmax 기반과 비교하여 우수한 성능을 보일 수 있다.
  • fp 샤프닝과 DWC를 도입하면 주의 랭크와 특징 다양성이 회복되어 정확도 상승을 가져오며(예: DeiT-T와 Swin-T 비교).
  • DeiT-Tiny, Swin-Tiny 및 기타 백본에서 FLatten은 비슷한 FLOPs와 매개변수로 더 높은 Top-1 정확도를 달성한다.
  • 추론 지연 분석은 CPU/GPU 하드웨어에서 기준선과 비교하여 최대 2.1배 빠른 실행 시간을 보이며 경쟁력 있는 정확도를 제공한다.
  • 벤치마크 전반에서 (ImageNet-1K, ADE20K, COCO) 유사한 계산 예산 하에 FLatten은 일관되게 베이스라인을 개선하거나 일치시킨다.
  • 네 가지 기존 선형 주의 설계와 비교하여 FLatten은 더 높은 정확도를 달성한다(예: DeiT-Tiny: 74.1 대 72.9–70.8; Swin-Tiny: 82.1 대 80.7–81.8).

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.