[논문 리뷰] Large Separable Kernel Attention: Rethinking the Large Kernel Attention Design in CNN
이 논문은 2D 깊이 지능형 컨볼루션을 연속된 1D 수평 및 수직 컨벌루션으로 분해하여, CNN 내에서 대규모 커널 컨볼루션의 계산 비용을 줄이는 대체 기법인 대규모 분리 커널 어텐션(LSKA)을 제안한다. LSKA는 LKA와 유사한 성능을 유지하면서도 FLOPs와 메모리 사용량을 최대 50% 감소시키며, 손상된 ImageNet 데이터에서 형태 편향과 강건성 또한 향상시킨다.
Visual Attention Networks (VAN) with Large Kernel Attention (LKA) modules have been shown to provide remarkable performance, that surpasses Vision Transformers (ViTs), on a range of vision-based tasks. However, the depth-wise convolutional layer in these LKA modules incurs a quadratic increase in the computational and memory footprints with increasing convolutional kernel size. To mitigate these problems and to enable the use of extremely large convolutional kernels in the attention modules of VAN, we propose a family of Large Separable Kernel Attention modules, termed LSKA. LSKA decomposes the 2D convolutional kernel of the depth-wise convolutional layer into cascaded horizontal and vertical 1-D kernels. In contrast to the standard LKA design, the proposed decomposition enables the direct use of the depth-wise convolutional layer with large kernels in the attention module, without requiring any extra blocks. We demonstrate that the proposed LSKA module in VAN can achieve comparable performance with the standard LKA module and incur lower computational complexity and memory footprints. We also find that the proposed LSKA design biases the VAN more toward the shape of the object than the texture with increasing kernel size. Additionally, we benchmark the robustness of the LKA and LSKA in VAN, ViTs, and the recent ConvNeXt on the five corrupted versions of the ImageNet dataset that are largely unexplored in the previous works. Our extensive experimental results show that the proposed LSKA module in VAN provides a significant reduction in computational complexity and memory footprints with increasing kernel size while outperforming ViTs, ConvNeXt, and providing similar performance compared to the LKA module in VAN on object recognition, object detection, semantic segmentation, and robustness tests.
연구 동기 및 목표
- 대규모 커널 깊이 지능형 컨볼루션으로 인한 FLOPs와 메모리 사용량의 제곱 증가를 해결하기 위해.
- 계산 복잡도를 증가시키지 않고도 매우 큰 커널을 CNN 어텐션 모듈에 적용할 수 있도록 하기 위해.
- 어텐션 메커니즘의 인덕티브 바이어스를 수정하여 모델의 강건성과 형태 편향을 향상시키기 위해.
- LSKA의 성능을 LKA, ViTs, ConvNeXt와 비교하여 손상된 ImageNet 및 후속 비전 작업에서 평가하기 위해.
- 시각 트랜스포머와 CNN에 대해 더 효율적이고 확장 가능한 어텐션 메커니즘을 제공하기 위해.
제안 방법
- LKA 모듈 내 2D 깊이 지능형 컨볼루션 커널을 연속된 1D 수평 및 수직 컨벌루션으로 분해하여 파라미터와 FLOP 증가를 줄이기.
- 동일한 어텐션 메커니즘을 사용하되, 분리 가능한 커널을 적용하여 1×1 컨벌루션을 통한 동일한 특징 정제 과정 유지.
- 이중 단계 설계를 적용: 먼저 작은 커널을 사용한 국소적 특징 추출, 그 다음 확장된 깊이 지능형 분리 컨벌루션을 통한 장거리 의존성 모델링.
- LSKA-trivial 및 LSKA 변형을 도입하여 커널 분해가 성능 및 효율성에 미치는 영향 평가.
- 모델 간 공정한 비교를 위해 LKA와 동일한 훈련 및 추론 파이프라인을 활용.
- 상호 정보를 통한 형태 및 텍스처 차원 분석을 통해 커널 크기와 분해에 의해 유도된 인덕티브 바이어스 이동 분석.
실험 결과
연구 질문
- RQ1분리 가능한 커널 분해가 성능을 희생시키지 않고도 CNN 내 대규모 커널 어텐션의 계산 비용을 줄일 수 있는가?
- RQ2LSKA 설계는 특징 표현에서 형태와 텍스처에 대한 모델의 인덕티브 바이어스에 어떤 영향을 미치는가?
- RQ3손상된 ImageNet 벤치마크에서 LSKA는 LKA, ViTs, ConvNeXt와 비교해 어떻게 강건성이 뛰어나게 되는가?
- RQ4LSKA에서 커널 크기를 증가시키는 것이 LKA와 비교해 효과적 수용 영역과 특징 표현에 어떤 영향을 미치는가?
- RQ5분리 가능한 커널 설계는 분류, 검출, 세그멘테이션과 같은 시각 작업에서 더 나은 일반화 및 효율성을 이끌어내는가?
주요 결과
- LSKA는 커널 크기가 65일 때 LKA 대비 FLOPs를 최대 50% 감소시키며, ImageNet에서 상위 1 정확도가 0.3% 감소하는 데 그친다.
- LSKA는 이미지 분류, 객체 검출, 세그멘테이션 작업 전반에서 LKA와 유사한 성능를 유지하면서도 메모리와 계산을 크게 줄였다.
- LSKA 설계는 특히 큰 커널을 사용할 경우 형태 표현에 대한 모델의 편향을 증가시키며, 형태-텍스처 차원 분석을 통해 이를 입증했다.
- LSKA는 모든 다섯 가지 손상된 ImageNet 벤치마크에서 ViTs와 ConvNeXt를 모두 앞서며 뛰어난 강건성을 입증했다.
- LSKA의 효과적 수용 영역은 커널 크기가 35, 53, 65일 때 포화 상태에 도달하며, LKA와 유사한 안정적인 장거리 모델링을 보였다.
- LSKA는 65×65 커널을 사용할 때 407만 개의 파라미터와 0.85 GFLOPs만으로 ImageNet에서 74.8%의 상위 1 정확도를 달성했으며, 동일한 커널을 사용하는 LKA는 473만 개의 파라미터와 1.12 GFLOPs를 필요로 했다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.