Skip to main content
QUICK REVIEW

[論文レビュー] Generalizable Industrial Visual Anomaly Detection with Self-Induction Vision Transformer

Haiming Yao, Yu, Wenyong|arXiv (Cornell University)|Nov 22, 2022
Anomaly Detection Techniques and Applications被引用数 9
ひとこと要約

本稿では、一般化可能でマルチカテゴリの産業用視覚異常検出を実現する自己誘導型ビジョン Transformer(SIVT)を提案する。事前学習済みの CNN を用いて局所的特徴を抽出し、ビジョン Transformer 内で補助誘導トークンを用いた新しい自己誘導メカニズムを導入することで特徴を再構築することで、SIVT は最先端の性能を達成し、Mvtec AD ベンチマークにおいて AUROC を 2.8–6.3 ポints、AP を 3.3–7.6 ポイント向上した。

ABSTRACT

Industrial vision anomaly detection plays a critical role in the advanced intelligent manufacturing process, while some limitations still need to be addressed under such a context. First, existing reconstruction-based methods struggle with the identity mapping of trivial shortcuts where the reconstruction error gap is legible between the normal and abnormal samples, leading to inferior detection capabilities. Then, the previous studies mainly concentrated on the convolutional neural network (CNN) models that capture the local semantics of objects and neglect the global context, also resulting in inferior performance. Moreover, existing studies follow the individual learning fashion where the detection models are only capable of one category of the product while the generalizable detection for multiple categories has not been explored. To tackle the above limitations, we proposed a self-induction vision Transformer(SIVT) for unsupervised generalizable multi-category industrial visual anomaly detection and localization. The proposed SIVT first extracts discriminatory features from pre-trained CNN as property descriptors. Then, the self-induction vision Transformer is proposed to reconstruct the extracted features in a self-supervisory fashion, where the auxiliary induction tokens are additionally introduced to induct the semantics of the original signal. Finally, the abnormal properties can be detected using the semantic feature residual difference. We experimented with the SIVT on existing Mvtec AD benchmarks, the results reveal that the proposed method can advance state-of-the-art detection performance with an improvement of 2.8-6.3 in AUROC, and 3.3-7.6 in AP.

研究の動機と目的

  • 再構築ベースの手法が産業用視覚異常検出において抱える限界、特に正常と異常サンプルの間で再構築誤差の差が小さいというアイデンティティマッピング問題を解決すること。
  • CNN の局所的受容 field の制限を克服し、ビジョン Transformer を用いてグローバルな文脈をモデル化すること。
  • 個々のカテゴリに特化した学習パラダイムを越えて、複数の製品カテゴリにわたる一般化可能な異常検出を実現すること。
  • 異常検出の特徴再構築の忠実度と異常局所化の正確性を向上させる自己教師付き再構築フレームワークを開発すること。
  • 補助トークンを用いた新しい自己誘導メカニズムを導入し、特徴再構築における自明なショートカットを防止すること。

提案手法

  • 入力画像から高意味的深層特徴マップを生成するために、事前学習済みの CNN を特徴抽出バックボーンとして採用する。
  • 抽出された特徴を補助誘導トークンを用いて再構築する自己誘導型ビジョン Transformer(SIVT)を提案する。
  • 補助誘導シーケンスを N 個(N=4)の互いに重複しない部分集合に分解し、特徴パッチ部分集合から補完的な意味を誘導する。
  • 計算コストを低減しつつ再構築精度を維持するために、パッチベースの特徴埋め込みを P=2 で実施する。
  • 異常ラベルを一切必要とせず、正常サンプルでの再構築損失のみを用いて自己教師付きでモデルを学習する。
  • 元の特徴と再構築特徴の間の意味的特徴残差差分を計算することで異常を検出する。

実験結果

リサーチクエスチョン

  • RQ1カテゴリ特化型モデルと比較して、ハイブリッド CNN-Transformer フレームワークはマルチカテゴリ産業用視覚異常検出における一般化性能を向上させ得るか?
  • RQ2ビジョン Transformer に補助誘導トークンを導入することで、異常検出における自明な再構築ショートカットが効果的に防止されるか?
  • RQ3誘導部分集合の数(N)は、自己誘導メカニズムの性能と効率にどのように影響を与えるか?
  • RQ4再構築精度と計算効率のバランスを最適化するには、特徴埋め込みのパッチサイズ(P)をどのように設定すべきか?
  • RQ5ヴァニラ Transformer に基づく再構築手法と比較して、自己誘導メカニズムは AUROC および AP をどの程度向上させるか?

主な発見

  • SIVT フレームワークは Mvtec AD ベンチマークで最先端の性能を達成し、先行手法と比較して AUROC を 2.8–6.3 ポイント、AP を 3.3–7.6 ポイント向上した。
  • 自己誘導メカニズムは異常サンプルの再構築誤差を顕著に低減し、ヴァニラ Transformer と比較して画像レベルの AUROC を 12.0、ピクセルレベルの AUROC を 6.8 向上させた。
  • 自己誘導メカニズムは、ヴァニラ Transformer ベースラインと比較して、画像レベルの AP を 6.2、ピクセルレベルの AP を 20.5 向上させた。
  • 誘導部分集合の数(N)を 4 に設定すると、性能と計算コストの最良のトレードオフが達成され、N=4 で性能がピークに達し、それ以上の N では性能が低下した。
  • パッチサイズ(P)を 2 に設定すると、再構築精度と効率の最適なバランスが得られ、P が 4 を超えると性能が急激に低下した。
  • SIVT モデルは複数の製品カテゴリにわたる優れた一般化性能を示し、各カテゴリごとの再トレーニングなしにクロスカテゴリ異常検出を可能にした。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。