[Paper Review] Provable More Data Hurt in High Dimensional Least Squares Estimator
This paper establishes the first finite-sample central limit theorem for prediction risk in high-dimensional least squares regression, proving that more data can degrade performance due to non-monotonic risk behavior in overparameterized regimes. Using random matrix theory, it derives confidence intervals that confirm the 'more data hurt' phenomenon when p/n → c > 1.
This paper investigates the finite-sample prediction risk of the high-dimensional least squares estimator. We derive the central limit theorem for the prediction risk when both the sample size and the number of features tend to infinity. Furthermore, the finite-sample distribution and the confidence interval of the prediction risk are provided. Our theoretical results demonstrate the sample-wise nonmonotonicity of the prediction risk and confirm "more data hurt" phenomenon.
Motivation & Objective
- To rigorously characterize the finite-sample distribution of prediction risk in high-dimensional least squares estimation.
- To resolve the paradox of 'more data hurt' in overparameterized models where increased sample size can worsen prediction performance.
- To provide a theoretical foundation for the non-monotonic behavior of prediction risk by analyzing second-order fluctuations.
- To derive asymptotic confidence intervals for prediction risk that enable statistical inference in high-dimensional settings.
Proposed method
- Applies random matrix theory, particularly the central limit theorem for linear spectral statistics, to analyze the prediction risk of least squares estimators.
- Derives the limiting distribution of the prediction risk under the asymptotic regime where both n and p grow with p/n → c.
- Introduces standardized test statistics T_n, T_{n,0}, T_{n,1}, T_{n,2}, T_{n,3} that converge to standard normal distributions under the CLT.
- Uses two types of prediction risk: conditional on design matrix and coefficient, and unconditional risk under model assumptions.
- Employs theoretical tools from Bai & Silverstein (2004) and Bai et al. (2007) to derive convergence rates and limiting variances.
- Validates theoretical results via Monte Carlo simulations with 1000 repetitions under various distributions (Gaussian, gamma, t-distribution).
Experimental results
Research questions
- RQ1Can the 'more data hurt' phenomenon in high-dimensional regression be formally proven using finite-sample distributions?
- RQ2How does the prediction risk behave as both sample size n and feature count p grow to infinity with p/n → c?
- RQ3What is the finite-sample distribution of the prediction risk in overparameterized linear models?
- RQ4Can confidence intervals for prediction risk be constructed that reflect the non-monotonic behavior of risk with increasing n?
- RQ5How do second-order fluctuations in prediction risk affect the monotonicity of model performance with respect to sample size?
Key findings
- The finite-sample distribution of the prediction risk converges to a normal distribution under the joint asymptotic regime n,p→∞ with p/n→c, enabling statistical inference.
- The confidence intervals for prediction risk exhibit non-monotonic behavior: for some n₁ < n₂ in the overparameterized regime (c>1), the upper bound at n₁ is below the lower bound at n₂, confirming 'more data hurt'.
- For c=1/2, the empirical coverage rate of 95% confidence intervals approaches 95% as p increases, validating the CLT approximation.
- For c=2, the finite-sample performance of the improved statistic T_{n,3} is better than T_{n,2}, with coverage rates approaching nominal levels as p grows.
- The prediction risk exhibits a sample-wise double descent curve, with a peak in risk around p/n≈1, followed by a decline, explaining why more data can hurt.
- The theoretical results confirm that the 'more data hurt' phenomenon arises from the interplay between overparameterization and second-order fluctuations in risk, not just asymptotic bias.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.