[論文レビュー] A ConvNet for the 2020s
本論文は ConvNeXt を提案する。これは純粋な ConvNet で、ResNet の設計思想を Transformer に近い性能へと現代化し、注意機構を用いずに ImageNet 精度で競争力を発揮し、COCO や ADE20K でも高い結果を示す。
The "Roaring 20s" of visual recognition began with the introduction of Vision Transformers (ViTs), which quickly superseded ConvNets as the state-of-the-art image classification model. A vanilla ViT, on the other hand, faces difficulties when applied to general computer vision tasks such as object detection and semantic segmentation. It is the hierarchical Transformers (e.g., Swin Transformers) that reintroduced several ConvNet priors, making Transformers practically viable as a generic vision backbone and demonstrating remarkable performance on a wide variety of vision tasks. However, the effectiveness of such hybrid approaches is still largely credited to the intrinsic superiority of Transformers, rather than the inherent inductive biases of convolutions. In this work, we reexamine the design spaces and test the limits of what a pure ConvNet can achieve. We gradually "modernize" a standard ResNet toward the design of a vision Transformer, and discover several key components that contribute to the performance difference along the way. The outcome of this exploration is a family of pure ConvNet models dubbed ConvNeXt. Constructed entirely from standard ConvNet modules, ConvNeXts compete favorably with Transformers in terms of accuracy and scalability, achieving 87.8% ImageNet top-1 accuracy and outperforming Swin Transformers on COCO detection and ADE20K segmentation, while maintaining the simplicity and efficiency of standard ConvNets.
研究の動機と目的
- Transformer に触発された近代化による ConvNet 設計の見直しを行い、Vision Transformer との性能ギャップを縮める。
- データ拡張を超える性能向上に寄与するアーキテクチャ上および訓練上の意思決定を特定する。
- 物体検出やセマンティックセグメンテーションなどの下流タスクで ConvNet の競争力を示す。
提案手法
- AdamW、Mixup、CutMix、RandAugment などの Transformer 風の訓練手法を用いた ResNet ベースラインから開始。
- マクロ設計を徐々に現代化(ステージの計算比、ステム、パッチ化したステム)し、Transformer に似た構造へ。
- FLOPs を抑えつつ幅を増やすため、ResNeXt 風のグループ化/深さ方向畳み込みを採用。
- 表現能力を高めつつ総 FLOPs を削減するため、 inverted bottleneck を取り入れる。
- 大きなカーネルの depthwise 畳み込み(7x7)を探索し、ブロック内の depthwise conv の配置を再考して Transformer ブロックの特性を模倣。
- マイクロ設計の微調整(活性化関数と正規化の選択、ダウンサンプリング戦略、LN と BN の比較など)を適用して ConvNet の性能を最大化。
- ImageNet-1K/22K、COCO(Mask R-CNN、Cascade Mask R-CNN)、ADE20K で評価し、転移性とスケーラビリティを示す。
実験結果
リサーチクエスチョン
- RQ1純粋な ConvNet をどこまで近代化すれば、階層的な Vision Transformer と同等の精度とスケーラビリティに匹敵できるか。
- RQ2アーキテクチャ上および訓練手法(マクロ、マイクロ、訓練のコツ)のうち、ConvNet の性能に対して Transformer と比較して最も影響を与えるのはどれか。
- RQ3Transformer 風のデータと手法で訓練された ConvNet バックボーンは、検出やセグメンテーションなどの下流タスクで Swin Transformer を上回ることができるか。
- RQ4大規模な事前学習(ImageNet-22K)が、ConvNet に対する Transformer の帰納的バイアスの利点を打ち消すか。
主な発見
- ConvNeXt 系は、サイズと解像度により ImageNet-1K の top-1 精度が約 82–87% に達する。
- ConvNeXt は、同等の FLOPs でいくつかの設定において Swin Transformer を上回る(例: 224^2 の ConvNeXt-B vs Swin-B: 83.8% vs 83.5%)。
- ConvNeXt-B の 384^2 では top-1 精度 85.1% に達し、同様の計算量下で Swin-B より高いスループット。
- ImageNet-22K 事前学習後、ConvNeXt-XL は 87.8% の top-1 を達成し、強力なスケーラビリティを示す。
- 下流タスクでは、ConvNeXt バックボーンは COCO の検出/セグメンテーションおよび ADE20K のセグメンテーションで Swin Transformers と同等かそれを上回り、しばしばより高いスループットを示す。
- Isotropic ConvNeXt blocks trained in ViT-like settings perform on par with ViT variants, showing competitive block design without non-convolutional attention.
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。