Skip to main content
QUICK REVIEW

[論文レビュー] The Shaped Transformer: Attention Models in the Infinite Depth-and-Width Limit

Lorenzo Noci, C. F. Li|arXiv (Cornell University)|Jun 30, 2023
Model Reduction and Neural NetworksPhysics and Astronomy被引用数 3
ひとこと要約

本稿では、深さと幅の無限極限において、Softmax出力を恒等写像を中心に、幅依存の温度パrameter τでロジットをスケーリングすることで、深く広いTransformerの安定化を図る「形状付きTransformer」を導入する。自己注意機構の変形により、共分散行列の確率的微分方程式(SDE)が導出され、残差接続とアーキテクチャの形状付けがランク崩壊を防ぎ、良好に条件付けられた表現を維持することを示す。シミュレーションにより、有限モデルにおけるSDEの正確性が確認された。

ABSTRACT

In deep learning theory, the covariance matrix of the representations serves as a proxy to examine the network's trainability. Motivated by the success of Transformers, we study the covariance matrix of a modified Softmax-based attention model with skip connections in the proportional limit of infinite-depth-and-width. We show that at initialization the limiting distribution can be described by a stochastic differential equation (SDE) indexed by the depth-to-width ratio. To achieve a well-defined stochastic limit, the Transformer's attention mechanism is modified by centering the Softmax output at identity, and scaling the Softmax logits by a width-dependent temperature parameter. We examine the stability of the network through the corresponding SDE, showing how the scale of both the drift and diffusion can be elegantly controlled with the aid of residual connections. The existence of a stable SDE implies that the covariance structure is well-behaved, even for very large depth and width, thus preventing the notorious issues of rank degeneracy in deep attention models. Finally, we show, through simulations, that the SDE provides a surprisingly good description of the corresponding finite-size model. We coin the name shaped Transformer for these architectural modifications.

研究の動機と目的

  • 深く広いTransformerの不安定性、特に初期化時のランク崩壊を解消すること。
  • 深さと幅が発散する際にも良好に動作する理論的裏付けに基づいた安定した注意メカニズムの開発。
  • 確率的微分方程式(SDE)を用いて、共分散構造の取り扱いやすい極限的記述を導出すること。
  • 理論的SDEモデルを有限サイズのニューラルネットワークのシミュレーションと照合して検証すること。
  • アーキテクチャの形状付け(恒等写像中心のSoftmaxと温度スケーリング)が、深さ付き注意モデルにおける安定した学習を可能にすることを示すこと。

提案手法

  • Softmax注意メカニズムを変更し、出力を恒等写像を中心に、幅依存の温度パrameter τでロジットをスケーリングする。
  • 比例的無限深さ・幅極限(d/n → γ > 0)において、共分散行列の極限的確率的微分方程式(SDE)を導出する。
  • 残差(スキップ)接続を用いて、SDEのドリフト項と拡散項を洗練的に制御し、安定性を確保する。
  • スキップ接続を備えた形状付きReLUフィードフォワードネットワークを含む、既存のSDEフレームワークを拡張する。
  • 有限ネットワークへの忠実度と確率的性質を保つために、比例極限(d, n → ∞ かつ d/n → γ)を用いる。
  • 有限サイズのネットワークのシミュレーションを用いてSDEモデルを検証し、SDEの予測と実測の共分散分布を比較する。
The Shaped Transformer: Attention Models in the Infinite Depth-and-Width Limit

実験結果

リサーチクエスチョン

  • RQ1比例的無限深さ・幅極限下で、変更された注意メカニズムは、深く広いTransformerにおけるランク崩壊を防げるか?
  • RQ2d/n → γ > 0 の比例極限において、注意層の共分散構造はどのように振る舞うか?
  • RQ3確率的微分方程式(SDE)は、このようなモデルの初期化時における共分散ダイナミクスを正確に記述できるか?
  • RQ4どのようなアーキテクチャ的修正がSDEを安定化させ、良好に条件付けられた表現を保証するか?
  • RQ5SDEは、有限サイズの実世界のTransformerの挙動をどの程度正確に予測できるか?

主な発見

  • 形状付きTransformerは、極めて深い深さと広い幅ですら、安定した相関分布が1未満に収束するなど、良好に条件付けられた共分散行列を維持することでランク崩壊を防ぐ。
  • 極限における共分散ダイナミクスは、深さ対幅比γでインデックス付けされた明確なSDEで記述され、残差接続によりドリフト項と拡散項を制御可能である。
  • シミュレーションにより、SDEが有限サイズのモデルの共分散構造を驚くほど正確に記述していることが確認され、カーネル密度推定値が実測データと密接に一致する。
  • 提案された温度スケーリングと恒等写像中心のSoftmaxは、飽和を軽減し、注意メカニズムを線形化することで、既知の学習不安定要因を緩和する。
  • GLUEにおける微調整実験では、形状付きTransformerはベースラインを上回り、特に深さd=24のアーキテクチャにおいてCOLAとRTEで優位である。COLAではF1スコアで最大0.211の向上を達成した。
  • エントロピーの崩壊(退化したSoftmax分布の兆候)は、形状付きTransformerでは大きな学習率下でも抑制されており、ベースラインモデルとは対照的である。
The Shaped Transformer: Attention Models in the Infinite Depth-and-Width Limit

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。