[논문 리뷰] A ConvNet for the 2020s
본 논문은 ConvNeXt를 개발하여 ResNet 설계 아이디어를 Transformer 유사한 성능으로 현대화한 순수 ConvNet으로, 주의 메커니즘 없이도 ImageNet에서 경쟁력 있는 정확도와 COCO 및 ADE20K에서 강력한 성능을 달성합니다.
The "Roaring 20s" of visual recognition began with the introduction of Vision Transformers (ViTs), which quickly superseded ConvNets as the state-of-the-art image classification model. A vanilla ViT, on the other hand, faces difficulties when applied to general computer vision tasks such as object detection and semantic segmentation. It is the hierarchical Transformers (e.g., Swin Transformers) that reintroduced several ConvNet priors, making Transformers practically viable as a generic vision backbone and demonstrating remarkable performance on a wide variety of vision tasks. However, the effectiveness of such hybrid approaches is still largely credited to the intrinsic superiority of Transformers, rather than the inherent inductive biases of convolutions. In this work, we reexamine the design spaces and test the limits of what a pure ConvNet can achieve. We gradually "modernize" a standard ResNet toward the design of a vision Transformer, and discover several key components that contribute to the performance difference along the way. The outcome of this exploration is a family of pure ConvNet models dubbed ConvNeXt. Constructed entirely from standard ConvNet modules, ConvNeXts compete favorably with Transformers in terms of accuracy and scalability, achieving 87.8% ImageNet top-1 accuracy and outperforming Swin Transformers on COCO detection and ADE20K segmentation, while maintaining the simplicity and efficiency of standard ConvNets.
연구 동기 및 목표
- Transformer에서 영감을 받은 현대화를 통해 ConvNet 설계를 재점검하고 Vision Transformer와의 격차를 좁힌다.
- 데이터 확장 외에도 성능 향상을 가져오는 아키텍처 및 학습 결정 요소를 식별한다.
- 객체 탐지 및 의미 분할과 같은 다운스트림 작업에서 ConvNet의 경쟁력을 입증한다.
제안 방법
- Transformer 유사 학습 트릭(AdamW, Mixup, CutMix, RandAugment 등)으로 학습된 ResNet 베이스라인에서 시작한다.
- Transformer 유사 구조를 향해 매크로 디자인(스테이지 계산 비율, 스템, 패치화 스템)을 점진적으로 현대화한다.
- 폭을 증가시키되 FLOPs를 제어하는 ResNeXt 스타일의 그룹/깊이별 합성곱을 채택한다.
- 전반적인 FLOPs를 줄이면서 표현 용량을 늘리기 위해 역-병목 구조를 도입한다.
- 대형 커널 깊이별 합성곱(7x7)과 블록 내 깊이별 합성곱의 재배치를 통해 Transformer 블록 특성을 모방한다.
- ConvNet 성능 극대화를 위한 마이크로 디자인 트릭(활성화 및 정규화 선택, 다운샘플링 전략, LN vs BN)을 적용한다.
- 전이 가능성과 확장성을 입증하기 위해 ImageNet-1K/22K, COCO(Mask R-CNN, Cascade Mask R-CNN), ADE20K에서 평가한다.
실험 결과
연구 질문
- RQ1순수 ConvNet를 얼마나 현대화해 계층형 Vision Transformer와 정확도 및 확장성에서 대등할 수 있는가?
- RQ2매크로, 마이크로 및 학습 트릭 등 어떤 아키텍처 및 학습 선택이 Transformer와 비교해 ConvNet 성능에 가장 큰 영향을 미치는가?
- RQ3Transformer 스타일 데이터와 트릭으로 학습된 ConvNet 백본이 다운스트림 작업에서 Swin Transformer를 능가할 수 있는가?
- RQ4대규모 사전 학습(이미지넷-22K) 은 ConvNet 대비 Transformer의 귀납적 편향 이점을 없애는가?
주요 결과
- ConvNeXt 변종은 크기와 해상도에 따라 ImageNet-1K 상위 1% 정확도(대략 82–87%)를 달성한다.
- 동일 FLOPs에서 여러 구성에서 Swin Transformer를 능가하는 ConvNeXt의 성능(예: ConvNeXt-B vs Swin-B 224^2에서 83.8% 대 83.5%)을 보인다.
- 384^2에서 ConvNeXt-B는 85.1%의 상위 1% 정확도에 도달하고 유사한 컴퓨트에서 Swin-B보다 처리량이 더 높다.
- ImageNet-22K 사전 학습 후 ConvNeXt-XL은 87.8% 상위 1% 정확도로 강력한 확장성을 보여준다.
- 다운스트림 작업에서 ConvNeXt 백본은 COCO 탐지/분할 및 ADE20K 분할에서 Swin Transformer과 일치하거나 능가하며 종종 더 나은 처리량을 보인다.
- ViT 유사 설정에서 학습된 등방성 ConvNeXt 블록은 ViT 변종과 대등한 성능을 보이며 비컨볼루션 어텐션 없이도 경쟁력 있는 블록 설계를 보여준다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.