Skip to main content
QUICK REVIEW

[論文レビュー] Random Fully Connected Neural Networks as Perturbatively Solvable Hierarchies

Boris Hanin|arXiv (Cornell University)|Apr 3, 2022
Neural Networks and Applications被引用数 4
ひとこと要約

本稿は、ガウス分布に従う重みとバイアスをもつランダムな全結合ニューラルネットワークの摂動的枠組みを構築し、出力とその微分の共同コマリヤントが $1/n$ のべき級数として解ける階層を示している。主な結果は、深さと幅の比 $L/n$(有効深さと呼ばれる)が非ガウス的フラクチュエーションと相関を支配し、勾配の爆発・消失問題を制御することであり、その影響は一次的に $L/n$ のスケールで現れる。

ABSTRACT

This article considers fully connected neural networks with Gaussian random weights and biases as well as $L$ hidden layers, each of width proportional to a large parameter $n$. For polynomially bounded non-linearities we give sharp estimates in powers of $1/n$ for the joint cumulants of the network output and its derivatives. Moreover, we show that network cumulants form a perturbatively solvable hierarchy in powers of $1/n$ in that $k$-th order cumulants in one layer have recursions that depend to leading order in $1/n$ only on $j$-th order cumulants at the previous layer with $j\leq k$. By solving a variety of such recursions, however, we find that the depth-to-width ratio $L/n$ plays the role of an effective network depth, controlling both the scale of fluctuations at individual neurons and the size of inter-neuron correlations. Thus, while the cumulant recursions we derive form a hierarchy in powers of $1/n$, contributions of order $1/n^k$ often grow like $L^k$ and are hence non-negligible at positive $L/n$. We use this to study a somewhat simplified version of the exploding and vanishing gradient problem, proving that this particular variant occurs if and only if $L/n$ is large. Several key ideas in this article were first developed at a physics level of rigor in a recent monograph of Daniel A. Roberts, Sho Yaida, and the author. This article not only makes these ideas mathematically precise but also significantly extends them, opening the way to obtaining corrections to all orders in $1/n$.

研究の動機と目的

  • 有限幅効果をランダムな全結合ニューラルネットワークで数学的に厳密に研究するための枠組みを提供すること。
  • 深さと幅が、ネットワーク出力および勾配の統計的性質に与える影響を同定すること。
  • コマリヤントに基づく解析を用いて、無限幅極限を超えた一般かつ非漸近的な設定で勾配の爆発・消失問題を解明すること。
  • 従来の物理的直感に基づく結果を、$1/n$ の完全な高次摂動的形式へと拡張すること。

提案手法

  • 層間をまたがる出力とその微分の共同コマリヤントの正確な再帰的関係を、$1/n$ のすべての次数で導出する。
  • $1/n$ の摂動展開を用いて、$k$ 階のコマリヤントが、前層の低階数コマリヤントにのみ依存することを示す。
  • $n, L \to \infty$ で $L/n \to \xi \in [0, \infty)$ となる二重スケーリング極限を導入し、正の $\xi$ で非ガウス的挙動が現れることを明らかにする。
  • 2次、3次、4次コマリヤントの明示的漸近展開を計算し、勾配の分散と相関構造を含む。
  • 形式的枠組みを用いて、入力およびパラメータに関して勾配分散の一次的挙動を導出し、$L/n$ のスケーリングを示す。
  • コマリヤント母関数とウィックの定理を用いて、$n \to \infty$ の極限における高次モーメントと相関を計算する。

実験結果

リサーチクエスチョン

  • RQ1有限幅効果は、深くランダムなニューラルネットワークにおける出力と勾配の連合分布にどのように現れるか?
  • RQ2深さと幅の比 $L/n$ は、ランダムネットワークにおけるフラクチュエーションと相関のスケールを決定づける役割を果たすか?
  • RQ3無限幅極限を超えて、ランダムな全結合ネットワークにおける勾配の爆発・消失問題を数学的に特徴づけられるか?
  • RQ4$1/n$ のコマリヤント再帰は、有限幅におけるガウス過程極限への体系的補正を可能にするか?

主な発見

  • 非線形関数が $K_* = 0$ の普遍性クラスに属する場合、ネットワーク出力の $2k$ 階コマリヤントは $k = 2,3,4$ に対して $(L/n)^{k/2 - 1}$ のように増大し、$L/n$ が非ガウス的性質を制御することを示している。
  • 出力の入力または1層目パラメータに関する勾配の分散は、$1/n$ の一次で $L/n$ のスケーリングを示し、勾配の爆発・消失問題の明確な数学的特徴づけが得られた。
  • 有効深さ $L/n$ は、ニューロン間相関と単一ニューロンのフラクチュエーションの両方を支配し、$1/n^k$ 階で $L^k$ のスケールで影響が増大するため、$L/n$ が支配的制御パrameterとなる。
  • コマリヤント階層は摂動的に解ける:層 $\ell+1$ の $k$ 階コマリヤントは、$1/n$ の一次で、層 $\ell$ の $j \leq k$ 階コマリヤントにのみ依存する。
  • 二重スケーリング極限 $n, L \to \infty$ で $L/n \to \xi$ の下で、$L/n \to 0$ の領域では捉えきれない非ガウス的かつ非線形効果が現れる。
  • 明示的漸近展開により $S_{(11)}^{(\ell)} \sim \frac{\ell}{3n} K_{(11)}^{(\ell)}$ が得られ、勾配分散が $L/n$ に対して線形にスケーリングすることを確認した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。