Skip to main content
QUICK REVIEW

[Paper Review] Provable More Data Hurt in High Dimensional Least Squares Estimator

Zeng Li, Chuanlong Xie|arXiv (Cornell University)|Aug 14, 2020
Random Matrices and Applications35 references6 citations
TL;DR

This paper establishes the first finite-sample central limit theorem for prediction risk in high-dimensional least squares regression, proving that more data can degrade performance due to non-monotonic risk behavior in overparameterized regimes. Using random matrix theory, it derives confidence intervals that confirm the 'more data hurt' phenomenon when p/n → c > 1.

ABSTRACT

This paper investigates the finite-sample prediction risk of the high-dimensional least squares estimator. We derive the central limit theorem for the prediction risk when both the sample size and the number of features tend to infinity. Furthermore, the finite-sample distribution and the confidence interval of the prediction risk are provided. Our theoretical results demonstrate the sample-wise nonmonotonicity of the prediction risk and confirm "more data hurt" phenomenon.

Motivation & Objective

  • To rigorously characterize the finite-sample distribution of prediction risk in high-dimensional least squares estimation.
  • To resolve the paradox of 'more data hurt' in overparameterized models where increased sample size can worsen prediction performance.
  • To provide a theoretical foundation for the non-monotonic behavior of prediction risk by analyzing second-order fluctuations.
  • To derive asymptotic confidence intervals for prediction risk that enable statistical inference in high-dimensional settings.

Proposed method

  • Applies random matrix theory, particularly the central limit theorem for linear spectral statistics, to analyze the prediction risk of least squares estimators.
  • Derives the limiting distribution of the prediction risk under the asymptotic regime where both n and p grow with p/n → c.
  • Introduces standardized test statistics T_n, T_{n,0}, T_{n,1}, T_{n,2}, T_{n,3} that converge to standard normal distributions under the CLT.
  • Uses two types of prediction risk: conditional on design matrix and coefficient, and unconditional risk under model assumptions.
  • Employs theoretical tools from Bai & Silverstein (2004) and Bai et al. (2007) to derive convergence rates and limiting variances.
  • Validates theoretical results via Monte Carlo simulations with 1000 repetitions under various distributions (Gaussian, gamma, t-distribution).

Experimental results

Research questions

  • RQ1Can the 'more data hurt' phenomenon in high-dimensional regression be formally proven using finite-sample distributions?
  • RQ2How does the prediction risk behave as both sample size n and feature count p grow to infinity with p/n → c?
  • RQ3What is the finite-sample distribution of the prediction risk in overparameterized linear models?
  • RQ4Can confidence intervals for prediction risk be constructed that reflect the non-monotonic behavior of risk with increasing n?
  • RQ5How do second-order fluctuations in prediction risk affect the monotonicity of model performance with respect to sample size?

Key findings

  • The finite-sample distribution of the prediction risk converges to a normal distribution under the joint asymptotic regime n,p→∞ with p/n→c, enabling statistical inference.
  • The confidence intervals for prediction risk exhibit non-monotonic behavior: for some n₁ < n₂ in the overparameterized regime (c>1), the upper bound at n₁ is below the lower bound at n₂, confirming 'more data hurt'.
  • For c=1/2, the empirical coverage rate of 95% confidence intervals approaches 95% as p increases, validating the CLT approximation.
  • For c=2, the finite-sample performance of the improved statistic T_{n,3} is better than T_{n,2}, with coverage rates approaching nominal levels as p grows.
  • The prediction risk exhibits a sample-wise double descent curve, with a peak in risk around p/n≈1, followed by a decline, explaining why more data can hurt.
  • The theoretical results confirm that the 'more data hurt' phenomenon arises from the interplay between overparameterization and second-order fluctuations in risk, not just asymptotic bias.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.