Skip to main content
QUICK REVIEW

[Paper Review] Exact Gap between Generalization Error and Uniform Convergence in Random Feature Models

Zitong Yang, Yu Bai|arXiv (Cornell University)|Mar 8, 2021
Stochastic Gradient Optimization Techniques49 references4 citations
TL;DR

This paper provides the first exact asymptotic analysis of the gap between generalization error and uniform convergence bounds in nonlinear random feature models. It derives precise expressions for classical uniform convergence, uniform convergence over interpolators, and the risk of the minimum norm interpolator, showing that while classical bounds can be vacuous in noisy settings, uniform convergence over interpolators remains non-vacuous and tightly controls test error.

ABSTRACT

Recent work showed that there could be a large gap between the classical uniform convergence bound and the actual test error of zero-training-error predictors (interpolators) such as deep neural networks. To better understand this gap, we study the uniform convergence in the nonlinear random feature model and perform a precise theoretical analysis on how uniform convergence depends on the sample size and the number of parameters. We derive and prove analytical expressions for three quantities in this model: 1) classical uniform convergence over norm balls, 2) uniform convergence over interpolators in the norm ball (recently proposed by Zhou et al. (2020)), and 3) the risk of minimum norm interpolator. We show that, in the setting where the classical uniform convergence bound is vacuous (diverges to $\infty$), uniform convergence over the interpolators still gives a non-trivial bound of the test error of interpolating solutions. We also showcase a different setting where classical uniform convergence bound is non-vacuous, but uniform convergence over interpolators can give an improved sample complexity guarantee. Our result provides a first exact comparison between the test errors and uniform convergence bounds for interpolators beyond simple linear models.

Motivation & Objective

  • To resolve the discrepancy between classical uniform convergence bounds and actual generalization errors in overparametrized models.
  • To analyze how uniform convergence behaves specifically over interpolating solutions in nonlinear random feature models.
  • To compare classical uniform convergence over norm balls with a refined version restricted to interpolators.
  • To quantify the exact asymptotic behavior of generalization error and uniform convergence bounds in high-dimensional settings.
  • To establish a precise theoretical framework for understanding generalization in the interpolating regime beyond linear models.

Proposed method

  • Derives exact asymptotic expressions for classical uniform convergence over norm balls of radius √A using high-dimensional random matrix theory.
  • Introduces and analyzes a refined uniform convergence bound over only interpolating functions within the same norm ball, as proposed by Zhou et al. (2020).
  • Computes the risk of the minimum norm interpolator using results from Mei & Montanari (2019) in the high-dimensional limit.
  • Uses Gegenbauer and Hermite polynomials to represent the activation function and analyze the spectral structure of the random feature kernel.
  • Establishes convergence of empirical quantities to their asymptotic counterparts under the limit n,d→∞ with ψ₁=n/d and ψ₂=N/d fixed.
  • Applies orthogonal decomposition in the space of spherical harmonics to derive exact expressions for the three key quantities: 𝒰, 𝒯, and 𝒓.

Experimental results

Research questions

  • RQ1How does classical uniform convergence behave in the presence of label noise when the model interpolates the training data?
  • RQ2Can uniform convergence over interpolators provide a non-vacuous bound on generalization error when classical uniform convergence becomes vacuous?
  • RQ3What is the exact relationship between the generalization error of the minimum norm interpolator and the uniform convergence bound over interpolators?
  • RQ4How does the sample complexity of generalization differ between the classical and interpolator-restricted uniform convergence frameworks?
  • RQ5In what regime does the refined uniform convergence over interpolators improve upon the classical bound in terms of sample complexity?

Key findings

  • In the noisy regime (τ² > 0), classical uniform convergence 𝒰 grows as √ψ₂ and becomes vacuous, while uniform convergence over interpolators 𝒯 remains bounded and converges to a constant.
  • The generalization error of the minimum norm interpolator 𝒓 decays as ψ₂⁻¹ in the noisy regime, while excess risk 𝒓 − τ² decays as ψ₂⁻¹.
  • In the noiseless regime (τ² = 0), the generalization error 𝒓 decays as ψ₂⁻², and the excess risk 𝒓 − τ² decays faster than ψ₂⁻¹.
  • Classical uniform convergence 𝒰 decays as ψ₂⁻¹/² in the noiseless regime, while uniform convergence over interpolators 𝒯 decays as ψ₂⁻¹/², showing improved sample complexity over classical bounds.
  • The gap between classical uniform convergence and actual generalization error is large in noisy settings, but the refined bound over interpolators remains tight and non-vacuous.
  • The asymptotic expressions for 𝒰, 𝒯, and 𝒓 are derived in terms of ψ₁ = N/d and ψ₂ = n/d, and are shown to concentrate in the high-dimensional limit.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.