Skip to main content
QUICK REVIEW

[論文レビュー] HiViT: Hierarchical Vision Transformer Meets Masked Image Modeling

Xiaosong Zhang, Yunjie Tian|arXiv (Cornell University)|May 30, 2022
Advanced Neural Network Applications被引用数 12
ひとこと要約

HiViTは、局所的なユニット間演算(例:シフトウィンドウ自己注意)を排除しながら、マルチスケール特徴マップを維持する階層的ビジョントランスフォーマーモデルを提案する。これにより、通常のViTと同様にトークンを逐次処理できるようになり、効率的なマスク画像モデリング(MIM)が可能となる。この設計により、MAE事前学習下でViT-Bを+0.6%上回るImageNet-1Kトップ-1精度を達成し、Swin-Bと比較して1.9倍高速な学習が実現した。

ABSTRACT

Recently, masked image modeling (MIM) has offered a new methodology of self-supervised pre-training of vision transformers. A key idea of efficient implementation is to discard the masked image patches (or tokens) throughout the target network (encoder), which requires the encoder to be a plain vision transformer (e.g., ViT), albeit hierarchical vision transformers (e.g., Swin Transformer) have potentially better properties in formulating vision inputs. In this paper, we offer a new design of hierarchical vision transformers named HiViT (short for Hierarchical ViT) that enjoys both high efficiency and good performance in MIM. The key is to remove the unnecessary "local inter-unit operations", deriving structurally simple hierarchical vision transformers in which mask-units can be serialized like plain vision transformers. For this purpose, we start with Swin Transformer and (i) set the masking unit size to be the token size in the main stage of Swin Transformer, (ii) switch off inter-unit self-attentions before the main stage, and (iii) eliminate all operations after the main stage. Empirical studies demonstrate the advantageous performance of HiViT in terms of fully-supervised, self-supervised, and transfer learning. In particular, in running MAE on ImageNet-1K, HiViT-B reports a +0.6% accuracy gain over ViT-B and a 1.9$ imes$ speed-up over Swin-B, and the performance gain generalizes to downstream tasks of detection and segmentation. Code will be made publicly available.

研究の動機と目的

  • 空間的に依存する局所的演算による階層的ビジョントランスフォーマーのMIMにおける非効率性を解消すること。
  • MIM効率性を阻害する『局所的ユニット間演算』を排除することで、階層的トランスフォーマーにおけるトークンの逐次処理を可能にすること。
  • 階層的特徴表現のパフォーマンス利点を維持しつつ、ViTと同等の計算の柔軟性を実現すること。
  • 自己教師あり、完全教師あり、および下流の検出・セグメンテーションタスクにおいて最先端の結果を示すこと。
  • 通常のViTと、Swin Transformerのような複雑な階層モデルの両者に対する、単純で効率的かつMIMに適した代替案を提供すること。

提案手法

  • 階層的ビジョントランスフォーマー内の演算を、ユニット内演算、グローバルユニット間演算、局所的ユニット間演算に分類する。
  • 『局所的ユニット間演算』(例:シフトウィンドウ自己注意、パッチマージング)がMIMにおけるトークン逐次処理を妨げる主な要因であると特定する。
  • Swin Transformerのメイン段階からすべての局所的ユニット間演算を削除し、グローバル自己注意と階層的特徴マップのみを保持する。
  • FLOPsとモデル容量を維持するために、初期段階の局所的演算を同等のユニット内MLPに置き換える。
  • 最終段階をメイン段階に統合することで、FLOPsを維持しつつアーキテクチャを単純化する。
  • マスクされたユニットを無視し、可視ユニットのみを処理することで、MIM事前学習中に可視トークンを完全に逐次処理可能にする。

実験結果

リサーチクエスチョン

  • RQ1階層的ビジョントランスフォーマーは、性能を損なわずに、通常のViTと同等のMIM効率性を達成できるか?
  • RQ2階層的トランスフォーマーにおける『局所的ユニット間演算』は認識精度に顕著な寄与をしているのか、それとも階層的構造自体が本質的な要因なのか?
  • RQ3Swin Transformerに対して最小限の変更(局所的自己注意の削除)を加えることで、ViTおよびSwinの両者を上回るMIM効率性と精度が得られるか?
  • RQ4得られたアーキテクチャ、HiViTは、自己教師あり、完全教師あり、および下流の検出・セグメンテーションタスクにおいて強力なパフォーマンスを維持できるか?
  • RQ5空間的に注意を意識するメカニズムを必要とせずに、MIMにおいて階層構造を効果的に活用できるか?

主な発見

  • HiViT-BはImageNet-1Kで83.8%のトップ-1精度を達成し、ViT-Bを+0.6%上回り、Swin-Bを+0.3%上回った。
  • MAE事前学習下で、HiViT-BはSwin-Bと比較して1.9倍高速に学習が可能であり、同時に精度でもViT-Bを上回った。
  • HiViT-Bの自己教師あり事前学習は、ImageNet-1Kで84.2%のトップ-1精度を達成し、教師ありベースラインと競合する水準だった。
  • 特徴ピラミッドの改善を加えることで、COCOオブジェクト検出とインスタンスセグメンテーションでそれぞれ51.2% AP^boxおよび44.2% AP^maskを達成した。
  • ADE20Kのセマンティックセグメンテーションでは、同じ事前学習プロトコル下でViT-Bを上回る51.2%のmIoUを達成した。
  • アブレーションスタディにより、局所的ユニット間演算の削除がパフォーマンスにほとんど影響を与えないことが確認され、コアな設計選択の妥当性が裏付けられた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。