Skip to main content
QUICK REVIEW

[Paper Review] The Eigenlearning Framework: A Conservation Law Perspective on Kernel Regression and Wide Neural Networks

James B. Simon, Madeline Dickens|arXiv (Cornell University)|Oct 8, 2021
Adversarial Robustness in Machine Learning4 citations
TL;DR

This paper introduces the Eigenlearning framework, a novel conservation law perspective that simplifies kernel ridge regression (KRR) generalization analysis by framing it as a competition among kernel eigenmodes for a fixed 'learnability' budget. The key contribution is a closed-form, interpretable theory for test risk and generalization metrics using only basic linear algebra, revealing deep connections to statistical physics and enabling new theoretical insights into deep learning generalization and adversarial robustness.

ABSTRACT

We derive simple closed-form estimates for the test risk and other generalization metrics of kernel ridge regression (KRR). Relative to prior work, our derivations are greatly simplified and our final expressions are more readily interpreted. These improvements are enabled by our identification of a sharp conservation law which limits the ability of KRR to learn any orthonormal basis of functions. Test risk and other objects of interest are expressed transparently in terms of our conserved quantity evaluated in the kernel eigenbasis. We use our improved framework to: i) provide a theoretical explanation for the "deep bootstrap" of Nakkiran et al (2020), ii) generalize a previous result regarding the hardness of the classic parity problem, iii) fashion a theoretical tool for the study of adversarial robustness, and iv) draw a tight analogy between KRR and a well-studied system in statistical physics.

Motivation & Objective

  • To simplify and reinterpret existing theoretical results on kernel ridge regression (KRR) generalization using a new conservation law.
  • To provide a transparent, closed-form framework for estimating test risk and other generalization metrics in KRR.
  • To explain empirical phenomena in deep learning, such as the 'deep bootstrap' and adversarial robustness, through a unified theoretical lens.
  • To draw a tight analogy between KRR and the free Fermi gas model in statistical physics, enabling cross-domain insights.
  • To offer a more accessible derivation of KRR generalization than prior work relying on replica theory or random matrix theory.

Proposed method

  • Introduces 'learnability' as a conserved quantity: the inner product between target and predicted functions, constrained by the number of training samples.
  • Derives a conservation law stating that total learnability across any orthonormal basis of functions cannot exceed the number of training samples, with equality at zero ridge parameter.
  • Expresses test risk, prediction covariance, and other metrics in terms of eigenmode-specific learnabilities using closed-form equations (e.g., Eq. 7–14).
  • Applies the framework to KRR in both discrete and continuous settings, with corrections for finite ridge and sample size via perturbation parameter δ/M.
  • Uses the framework to derive estimators for adversarial robustness via predicted function smoothness and to analyze the deep bootstrap phenomenon.
  • Establishes a formal analogy between KRR and the free Fermi gas, leveraging known physics results to inform KRR behavior.

Experimental results

Research questions

  • RQ1How can generalization metrics in kernel ridge regression be expressed in a closed-form, interpretable way using only basic linear algebra?
  • RQ2What is the theoretical basis for the observed 'deep bootstrap' phenomenon in wide neural networks, and how can it be explained via eigenmode competition?
  • RQ3How does the conservation of learnability constrain the ability of KRR to learn orthogonal target functions?
  • RQ4Can the framework provide a new theoretical tool for analyzing adversarial robustness in machine learning models?
  • RQ5What is the precise mathematical analogy between kernel ridge regression and the free Fermi gas in statistical physics?

Key findings

  • The total learnability across any orthonormal basis of functions is conserved and bounded by the number of training samples, with equality when the ridge parameter is zero.
  • Test risk and prediction covariance in KRR are expressed in closed-form using eigenmode learnabilities, offering a significant simplification over prior replica-theory-based derivations.
  • The framework explains the 'deep bootstrap' phenomenon by identifying two distinct training regimes—early and late—where different eigenmodes dominate learning.
  • A new estimator for predicted function smoothness is derived, providing a theoretical tool for studying adversarial robustness.
  • The system exhibits a tight analogy to the free Fermi gas, with eigenmode learnabilities mirroring fermion occupation numbers, enabling transfer of insights from statistical physics.
  • Analytical bounds on the ridge parameter κ are derived, showing it decreases with training sample size and is sensitive to the spectrum of kernel eigenvalues.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.