Skip to main content
QUICK REVIEW

[Paper Review] Global Convergence of Three-layer Neural Networks in the Mean Field Regime

Huy Tuan Pham, Phan-Minh Nguyen|arXiv (Cornell University)|May 11, 2021
Stochastic Gradient Optimization Techniques29 references4 citations
TL;DR

This paper establishes global convergence for unregularized three-layer feedforward neural networks under stochastic gradient descent in the mean field regime. By introducing a novel neuronal embedding framework that handles intertwined symmetry groups, the authors prove convergence to global optima without relying on convexity, leveraging a finite-time universal approximation property via algebraic topology arguments.

ABSTRACT

In the mean field regime, neural networks are appropriately scaled so that as the width tends to infinity, the learning dynamics tends to a nonlinear and nontrivial dynamical limit, known as the mean field limit. This lends a way to study large-width neural networks via analyzing the mean field limit. Recent works have successfully applied such analysis to two-layer networks and provided global convergence guarantees. The extension to multilayer ones however has been a highly challenging puzzle, and little is known about the optimization efficiency in the mean field regime when there are more than two layers. In this work, we prove a global convergence result for unregularized feedforward three-layer networks in the mean field regime. We first develop a rigorous framework to establish the mean field limit of three-layer networks under stochastic gradient descent training. To that end, we propose the idea of a extit{neuronal embedding}, which comprises of a fixed probability space that encapsulates neural networks of arbitrary sizes. The identified mean field limit is then used to prove a global convergence guarantee under suitable regularity and convergence mode assumptions, which -- unlike previous works on two-layer networks -- does not rely critically on convexity. Underlying the result is a universal approximation property, natural of neural networks, which importantly is shown to hold at extit{any} finite training time (not necessarily at convergence) via an algebraic topology argument.

Motivation & Objective

  • To establish global convergence of three-layer neural networks in the mean field regime, where optimization dynamics are governed by a nonlinear limit.
  • To address the conceptual challenge of intertwined symmetry groups in multilayer networks, which hinders prior mean field analyses.
  • To develop a rigorous framework for the mean field limit of three-layer networks trained via stochastic gradient descent.
  • To prove global convergence without relying on convexity, unlike prior two-layer analyses.
  • To demonstrate a universal approximation property that holds at any finite training time, not just at convergence.

Proposed method

  • Introduces the neuronal embedding—a fixed probability space that encapsulates neural networks of arbitrary widths, enabling analysis of the mean field limit.
  • Uses a measure-theoretic framework inspired by Sznitman (1991) and Mei et al. (2018) to rigorously connect the finite-width network to its mean field limit.
  • Establishes quantitative approximation bounds: the mean field limit approximates the network when $n_{\min}^{-1}\log n_{\max} \ll 1$, independent of data dimension.
  • Employs an algebraic topology argument to prove a universal approximation property valid at any finite time during training.
  • Proves convergence of the mean field limit to the global optimum under regularity and convergence mode assumptions, without assuming convexity.
  • Adapts techniques from Chizat & Bach (2018) but extends them to handle the non-convex, three-layer setting via the new framework.

Experimental results

Research questions

  • RQ1Can global convergence be established for three-layer neural networks in the mean field regime, despite the absence of convexity?
  • RQ2How can the intertwined symmetry groups in multilayer networks be formally addressed in the mean field limit?
  • RQ3Does a universal approximation property hold for three-layer networks at finite training times, not just at convergence?
  • RQ4What conditions ensure that the mean field limit accurately approximates the finite-width network during training?
  • RQ5Can the convergence guarantee be proven without relying on convexity, as in prior two-layer analyses?

Key findings

  • The neuronal embedding framework successfully captures the mean field limit of three-layer networks under SGD, even when multiple symmetry groups act simultaneously.
  • The mean field limit provides a good approximation of the finite-width network when $n_{\min}^{-1}\log n_{\max} \ll 1$, independent of data dimension.
  • Global convergence to the global optimum is achieved for the mean field limit under suitable regularity and convergence mode assumptions.
  • The proof does not rely on convexity, marking a conceptual advancement over prior two-layer mean field analyses.
  • A universal approximation property holds at any finite training time, established via an algebraic topology argument, which is crucial for the convergence result.
  • The framework generalizes beyond i.i.d. initialization and handles non-i.i.d. schemes, overcoming limitations of prior formulations.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.