Skip to main content
QUICK REVIEW

[Paper Review] Fundamental limits of overparametrized shallow neural networks for supervised learning

Francesco Camilli, Daria Tieplova|arXiv (Cornell University)|Jul 11, 2023
Neural Networks and ApplicationsComputer Science3 citations
TL;DR

This paper establishes rigorous information-theoretic bounds on the generalization performance of overparametrized two-layer neural networks trained on data generated by a teacher network with the same architecture. Using tools from spin glass theory and Gaussian equivalence principles, it proves that in a high-dimensional regime with sufficient training samples, the neural network behaves equivalently to a generalized linear model, achieving the same Bayes-optimal generalization error and mutual information bounds.

ABSTRACT

We carry out an information-theoretical analysis of a two-layer neural network trained from input-output pairs generated by a teacher network with matching architecture, in overparametrized regimes. Our results come in the form of bounds relating i) the mutual information between training data and network weights, or ii) the Bayes-optimal generalization error, to the same quantities but for a simpler (generalized) linear model for which explicit expressions are rigorously known. Our bounds, which are expressed in terms of the number of training samples, input dimension and number of hidden units, thus yield fundamental performance limits for any neural network (and actually any learning procedure) trained from limited data generated according to our two-layer teacher neural network model. The proof relies on rigorous tools from spin glasses and is guided by ``Gaussian equivalence principles'' lying at the core of numerous recent analyses of neural networks. With respect to the existing literature, which is either non-rigorous or restricted to the case of the learning of the readout weights only, our results are information-theoretic (i.e. are not specific to any learning algorithm) and, importantly, cover a setting where all the network parameters are trained.

Motivation & Objective

  • To understand the fundamental performance limits of overparametrized shallow neural networks in supervised learning.
  • To analyze how architecture and data availability constrain generalization error and mutual information in a teacher-student framework.
  • To rigorously establish the validity of Gaussian equivalence principles (GEPs) in information-theoretic settings where all network weights are trained.
  • To derive bounds on mutual information and generalization error that are independent of specific learning algorithms.
  • To bridge the gap between non-rigorous heuristic analyses and rigorous mathematical physics in overparametrized neural network regimes.

Proposed method

  • Employs a Bayesian-optimal teacher-student setup with random i.i.d. inputs and outputs generated by a two-layer teacher network.
  • Applies rigorous tools from spin glass theory to analyze the high-dimensional limit of the neural network's generalization performance.
  • Uses Gaussian equivalence principles (GEPs) to show that non-linear activations in overparametrized networks behave like linear models under specific scaling.
  • Derives explicit bounds on mutual information between training data and network weights by comparing the neural network to a simpler generalized linear model.
  • Performs Gaussian integration by parts and orthogonalization techniques to control higher-order corrections in the asymptotic expansion.
  • Establishes equivalence between the neural network and a generalized linear model in terms of optimal generalization error and mutual information, under the condition that the number of training samples scales appropriately with input dimension and number of hidden units.

Experimental results

Research questions

  • RQ1What are the fundamental information-theoretic limits of overparametrized two-layer neural networks in supervised learning?
  • RQ2Under what scaling regimes do non-linear neural networks behave like generalized linear models in terms of generalization performance?
  • RQ3How does the mutual information between training data and network weights relate to that of a simpler linear model in the overparametrized regime?
  • RQ4Can Gaussian equivalence principles be rigorously justified in settings where all network weights are trained, not just the readout layer?
  • RQ5What is the optimal generalization error achievable by any learning procedure in this overparametrized teacher-student setup?

Key findings

  • In the overparametrized regime with sufficient training samples, the mutual information between training data and network weights in a two-layer neural network is bounded by that of a generalized linear model.
  • The Bayes-optimal generalization error of the neural network is information-theoretically equivalent to that of a generalized linear model under the same scaling conditions.
  • The Gaussian equivalence principle (GEP) is rigorously validated in this setting, showing that non-linear activations lead to linear-like behavior in high-dimensional limits.
  • The bounds depend explicitly on the number of training samples, input dimension, and number of hidden units, providing a quantitative characterization of generalization limits.
  • The analysis holds for any learning algorithm, as it is information-theoretic and not algorithm-specific, establishing fundamental performance ceilings.
  • Higher-order corrections in the asymptotic expansion are shown to be negligible under the derived scaling regime, confirming the robustness of the equivalence to generalized linear models.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.