[論文レビュー] Scale Alone Does not Improve Mechanistic Interpretability in Vision Models
本研究では、人間参加者を用いた大規模な心理物理学実験を通じて、データセットおよびアーキテクチャのスケーリングが機械的解釈可能性を向上させるかを調査している。スケーリングによる解釈可能性の向上は見られず、現代のモデルは2014年のGoogLeNetと同等にしか解釈されない。性能向上にもかかわらず解釈可能性が低下している可能性を示唆し、解釈可能性のための明示的設計と自動評価メトリクスの開発を要請している。
In light of the recent widespread adoption of AI systems, understanding the internal information processing of neural networks has become increasingly critical. Most recently, machine vision has seen remarkable progress by scaling neural networks to unprecedented levels in dataset and model size. We here ask whether this extraordinary increase in scale also positively impacts the field of mechanistic interpretability. In other words, has our understanding of the inner workings of scaled neural networks improved as well? We use a psychophysical paradigm to quantify one form of mechanistic interpretability for a diverse suite of nine models and find no scaling effect for interpretability - neither for model nor dataset size. Specifically, none of the investigated state-of-the-art models are easier to interpret than the GoogLeNet model from almost a decade ago. Latest-generation vision models appear even less interpretable than older architectures, hinting at a regression rather than improvement, with modern models sacrificing interpretability for accuracy. These results highlight the need for models explicitly designed to be mechanistically interpretable and the need for more helpful interpretability methods to increase our understanding of networks at an atomic level. We release a dataset containing more than 130'000 human responses from our psychophysical evaluation of 767 units across nine models. This dataset facilitates research on automated instead of human-based interpretability evaluations, which can ultimately be leveraged to directly optimize the mechanistic interpretability of models.
研究の動機と目的
- データセットおよびアーキテクチャのスケーリングが、視覚モデルの機械的解釈可能性を向上させるかどうかを評価すること。
- 最近の大規模視覚モデル(例:ViT、ConvNeXt)が、GoogLeNetのような過去のアーキテクチャよりも解釈可能かどうかを調査すること。
- 自然例と特徴可視化という2つの標準的手法が、多様なモデル群においてどれほど効果的かを評価すること。
- 自動解釈可能性評価とモデル最適化を可能にするために、人間がアノテートした大規模な解釈可能性反応データセットを公開すること。
- モデルスケールが解釈可能性を必然的に向上させるという仮定に疑問を呈し、解釈可能性のための明示的設計を提唱すること。
提案手法
- 9つの最先端視覚モデルにおける767ユニットで、12万件のヒトによる反応を含む大規模な心理物理学実験を実施した。
- 2つの標準的手法を用いた:ImageNetからの高活性化画像(自然例)の特定と、勾配上昇を用いた特徴可視化の生成。
- Amazon Mechanical Turkの制御されたHITインターフェースを用いて、参加者がユニットが反応する特徴を特定できたかどうかのヒト判断を収集した。
- モデルタイプ、深さ、スケール(例:ViT対CNN、ResNet対ConvNeXt)の違いに応じた解釈可能性スコアの比較を通じて、解釈可能性を評価した。
- ピアソン相関を用いて、活性化スパarsityと解釈可能性の間の相関関係をテストし、予測的関係の有無を分析した。
- 全データセット(ImageNet Mechanistic Interpretability)を公開し、将来の自動解釈可能性評価およびモデル最適化を支援した。
実験結果
リサーチクエスチョン
- RQ1モデルやデータセットのスケーリングを拡大することで、視覚モデルの機械的解釈可能性が向上するか?
- RQ2最近の視覚モデル(例:ViT、ConvNeXt)は、GoogLeNetのような古いモデルよりも解釈可能か?
- RQ3標準的手法(自然例と特徴可視化)は、より大きなモデルにおいてより良い解釈結果をもたらすか?
- RQ4活性化スパarsityとヒトが感じる解釈可能性の間には相関があるか?
- RQ5人間がアノテートしたデータセットを用いて、将来のモデル開発のための自動解釈可能性メトリクスを学習できるか?
主な発見
- 顕著なスケーリング効果は認められなかった:9つの現代モデルのいずれも、約10年前のGoogLeNetよりも解釈されやすくなかった。
- ViT や ConvNeXt といった現代モデルは、ResNet-50 や GoogLeNet といった過去のCNNよりも低い解釈可能性スコアを示した。
- 活性化スパarsityと解釈可能性の間には弱く有意でない相関関係(r = 0.10, p = 0.39)が認められ、スパarsityが解釈可能性を予測できないことが示唆された。
- 構造的・スケール的に大きく異なるモデル間で解釈可能性に有意な差がなく、スケーリングが解釈可能性を向上させるという仮定に疑問を呈した。
- 12万件を超えるヒトによる反応を含む人間アノテートデータセットが公開され、将来の自動解釈可能性評価手法の開発を支援する。
- スケーリングに伴い解釈可能性が低下している可能性を示唆し、現代のモデルは正確性の向上の代償として解釈可能性を犠牲にしている可能性があり、解釈可能性のための明示的設計が不可欠であると考えられる。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。