Skip to main content
QUICK REVIEW

[Paper Review] Random Fully Connected Neural Networks as Perturbatively Solvable Hierarchies

Boris Hanin|arXiv (Cornell University)|Apr 3, 2022
Neural Networks and Applications4 citations
TL;DR

This paper develops a perturbative framework for analyzing random fully connected neural networks with Gaussian weights and biases, showing that joint cumulants of outputs and their derivatives form a hierarchy solvable in powers of $1/n$, where $n$ is width. The key result is that the depth-to-width ratio $L/n$, termed the effective depth, governs non-Gaussian fluctuations and correlations, and controls the exploding/vanishing gradient problem, with effects scaling as $L/n$ at leading order.

ABSTRACT

This article considers fully connected neural networks with Gaussian random weights and biases as well as $L$ hidden layers, each of width proportional to a large parameter $n$. For polynomially bounded non-linearities we give sharp estimates in powers of $1/n$ for the joint cumulants of the network output and its derivatives. Moreover, we show that network cumulants form a perturbatively solvable hierarchy in powers of $1/n$ in that $k$-th order cumulants in one layer have recursions that depend to leading order in $1/n$ only on $j$-th order cumulants at the previous layer with $j\leq k$. By solving a variety of such recursions, however, we find that the depth-to-width ratio $L/n$ plays the role of an effective network depth, controlling both the scale of fluctuations at individual neurons and the size of inter-neuron correlations. Thus, while the cumulant recursions we derive form a hierarchy in powers of $1/n$, contributions of order $1/n^k$ often grow like $L^k$ and are hence non-negligible at positive $L/n$. We use this to study a somewhat simplified version of the exploding and vanishing gradient problem, proving that this particular variant occurs if and only if $L/n$ is large. Several key ideas in this article were first developed at a physics level of rigor in a recent monograph of Daniel A. Roberts, Sho Yaida, and the author. This article not only makes these ideas mathematically precise but also significantly extends them, opening the way to obtaining corrections to all orders in $1/n$.

Motivation & Objective

  • To provide a mathematically rigorous framework for studying finite-width effects in random fully connected neural networks.
  • To characterize how depth and width jointly influence the statistical properties of network outputs and gradients.
  • To resolve the exploding and vanishing gradient problem in a general, non-asymptotic setting using cumulant-based analysis.
  • To extend prior physics-inspired results into a fully rigorous, higher-order perturbative formalism in $1/n$.

Proposed method

  • Derives exact recursive relations for joint cumulants of network outputs and their derivatives across layers, valid to all orders in $1/n$.
  • Uses a perturbative expansion in $1/n$ to show that $k$-th order cumulants depend only on lower-order cumulants from the previous layer at leading order.
  • Introduces a double scaling limit where $n, L \to \infty$ with $L/n \to \xi \in [0, \infty)$, revealing non-Gaussian behavior at positive $\xi$.
  • Computes explicit asymptotics for second, third, and fourth-order cumulants, including variance and correlation structure of gradients.
  • Applies the formalism to derive the leading-order behavior of gradient variance with respect to input and parameters, showing scaling as $L/n$.
  • Uses cumulant generating functions and Wick's theorem to compute higher-order moments and correlations in the large-$n$ limit.

Experimental results

Research questions

  • RQ1How do finite-width effects manifest in the joint distribution of outputs and gradients in deep random neural networks?
  • RQ2What is the role of the depth-to-width ratio $L/n$ in determining the scale of fluctuations and correlations in random networks?
  • RQ3Can the exploding and vanishing gradient problem be characterized mathematically in random fully connected networks beyond the infinite-width limit?
  • RQ4How do cumulant recursions in $1/n$ enable systematic corrections to the Gaussian process limit at finite width?

Key findings

  • For non-linearities in the $K_* = 0$ universality class, the $2k$-th cumulants of the network output grow as $(L/n)^{k/2 - 1}$ for $k = 2,3,4$, indicating that $L/n$ controls non-Gaussianity.
  • The variance of the gradient of the network output with respect to inputs or first-layer parameters scales as $L/n$ at leading order in $1/n$, providing a precise mathematical characterization of the exploding/vanishing gradient problem.
  • The effective depth $L/n$ governs both inter-neuron correlations and single-neuron fluctuations, with effects growing like $L^k$ at order $1/n^k$, making $L/n$ the dominant control parameter.
  • The cumulant hierarchy is perturbatively solvable: $k$-th order cumulants at layer $\ell+1$ depend only on $j \leq k$-th order cumulants at layer $\ell$ at leading order in $1/n$.
  • In the double scaling limit $n, L \to \infty$ with $L/n \to \xi$, the network exhibits non-Gaussian and non-linear effects not captured by the $L/n \to 0$ regime.
  • Explicit asymptotics show $S_{(11)}^{(\ell)} \sim \frac{\ell}{3n} K_{(11)}^{(\ell)}$, confirming that gradient variance scales linearly with $L/n$.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.