[Paper Review] Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
The paper proves dimension-free non-asymptotic bounds for mean-field approximations of SGD in two-layer networks, extends to unbounded activations and noisy SGD, and connects the mean-field dynamics to kernel ridge regression in a kernel limit.
We consider learning two layer neural networks using stochastic gradient descent. The mean-field description of this learning dynamics approximates the evolution of the network weights by an evolution in the space of probability distributions in $R^D$ (where $D$ is the number of parameters associated to each neuron). This evolution can be defined through a partial differential equation or, equivalently, as the gradient flow in the Wasserstein space of probability distributions. Earlier work shows that (under some regularity assumptions), the mean field description is accurate as soon as the number of hidden units is much larger than the dimension $D$. In this paper we establish stronger and more general approximation guarantees. First of all, we show that the number of hidden units only needs to be larger than a quantity dependent on the regularity properties of the data, and independent of the dimensions. Next, we generalize this analysis to the case of unbounded activation functions, which was not covered by earlier bounds. We extend our results to noisy stochastic gradient descent. Finally, we show that kernel ridge regression can be recovered as a special limit of the mean field analysis.
Motivation & Objective
- Motivate and analyze the mean-field description of learning in two-layer neural networks under SGD.
- Derive dimension-free non-asymptotic approximation guarantees between SGD and the PDE/mean-field dynamics.
- Extend the analysis to unbounded activations and noisy SGD.
- Show how kernel ridge regression emerges as a kernel limit of the mean-field dynamics.
Proposed method
- Model the network as an average over N neurons with parameters θi=(ai,wi) and activation σ*, and study the empirical distribution ^(N) of neurons.
- Formulate the mean-field evolution as a PDE on the space of distributions ρt with Ψ and its components V and U.
- Prove dimension-free bounds showing SGD approximates the mean-field PDE with error decaying as 1/√N and terms in √(D+log N) and √ε.
- Extend to noisy SGD leading to a diffusion-DD PDE, and provide bounds under strengthened assumptions.
- Introduce a kernel-limit by a scale α, yielding a residual dynamics that aligns with kernel ridge regression in the short-time/linearized regime.
- Demonstrate a coupled dynamics between residuals and kernel evolution, and analyze the kernel limit via linearized dynamics.
Experimental results
Research questions
- RQ1Under what conditions does the mean-field PDE provide a dimension-free approximation to SGD for two-layer networks?
- RQ2How do unbounded activations and noisy SGD affect the accuracy of the mean-field approximation?
- RQ3Can kernel ridge regression be recovered as a kernel limit of the mean-field dynamics, and what is the nature of this limit?
- RQ4What changes when introducing a scale parameter α in the kernel/mean-field coupling, and how does it affect convergence and residual dynamics?
- RQ5What are the quantitative rates and dependencies (on N, D, ε, T) for the approximation bounds between SGD and the mean-field description?
Key findings
- The number of hidden units N needs to exceed a data-regularity-dependent quantity, independent of the dimension D, for the mean-field approximation to hold.
- Dimension-free bounds are established for both bounded and unbounded activations under suitable conditions.
- Noisy SGD admits a dimension-free bound in the fixed-coefficient setting, with a diffusion term in the PDE, though some cases with unbounded coefficients lose full dimension-free scaling.
- Kernel ridge regression can be recovered as a special limit of the mean-field analysis via a short-time, linearized dynamics.
- A kernel-limit dynamics coupled to residual evolution shows a time-varying data-dependent kernel, providing a bridge between mean-field SGD and kernel methods.
- The results generalize prior work by relaxing activation bounds, incorporating noise, and proving dimension-free dependence.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.