Skip to main content
QUICK REVIEW

[Paper Review] Asymptotics of Wide Networks from Feynman Diagrams

Ethan Dyer, Guy Gur-Ari|arXiv (Cornell University)|Sep 25, 2019
advanced mathematical theories36 references45 citations
TL;DR

This paper develops a general, diagrammatic method (inspired by Feynman diagrams) to bound the asymptotics of correlation functions for wide neural networks and applies it to training dynamics, yielding finite-width corrections and tighter SGD/NTK results.

ABSTRACT

Understanding the asymptotic behavior of wide networks is of considerable interest. In this work, we present a general method for analyzing this large width behavior. The method is an adaptation of Feynman diagrams, a standard tool for computing multivariate Gaussian integrals. We apply our method to study training dynamics, improving existing bounds and deriving new results on wide network evolution during stochastic gradient descent. Going beyond the strict large width limit, we present closed-form expressions for higher-order terms governing wide network training, and test these predictions empirically.

Motivation & Objective

  • Motivate understanding of wide (large-width) neural networks and their training behavior.
  • Introduce a general method to bound asymptotics of network correlation functions using a diagrammatic approach.
  • Apply the method to training dynamics to tighten bounds on gradient flow and SGD evolution.
  • Provide finite-width corrections to the infinite-width limit and connect to Neural Tangent Kernel (NTK) dynamics.

Proposed method

  • Define correlation functions as ensemble averages of network outputs and derivatives with respect to parameters.
  • Adapt Feynman diagram techniques to bound large-width behavior via a conjectured scaling bound (Conjecture 1).
  • Prove the conjecture exactly for deep linear networks, and provide evidence/partial proofs for networks with nonlinearity (ReLU, tanh) and non-Gaussian initializations.
  • Derive finite-width corrections to training dynamics by expanding around the infinite-width limit and solving coupled equations for the kernel and network map.
  • Develop and employ Feynman rules to compute diagrammatic contributions and bound their n-dependence.
  • Show that SGD updates are linear in the learning rate in the large-width limit and compute leading finite-width corrections to the NTK and network evolution.

Experimental results

Research questions

  • RQ1Can correlation functions of wide networks be bounded in the large-width limit using a diagrammatic method?
  • RQ2How do finite-width corrections modify the Neural Tangent Kernel and network evolution under gradient flow and SGD?
  • RQ3Do the proposed bounds extend beyond deep linear networks to networks with nonlinear activations like ReLU and tanh?
  • RQ4What are the leading-order finite-width corrections to training dynamics and spectral properties of the NTK and Hessian?

Key findings

  • A general conjecture bounds correlation functions with a width-dependent exponent determined by cluster components (even/odd) of the contraction graph.
  • For deep linear networks, the conjecture is proven; for networks with ReLU or one hidden layer with smooth activation, the conjecture holds under certain conditions.
  • In training dynamics, the NTK is constant up to O(n^{-1}) corrections under both gradient flow and SGD.
  • Leading finite-width corrections to the network map and NTK are derived in closed form, expressed via O_s functions and integrals over the NTK spectrum.
  • Empirical experiments (e.g., two-class MNIST) validate O(n^{-1}) scaling and the predicted finite-width corrections across activations and initialization schemes.
  • The method yields tighter bounds on kernel evolution than prior results and shows linear-in-learning-rate behavior in SGD in the large-width regime.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.