Skip to main content
QUICK REVIEW

[Paper Review] Early Stopping for Nonparametric Testing

Meimei Liu, Guang Cheng|arXiv (Cornell University)|May 25, 2018
Statistical Methods and Inference14 references3 citations
TL;DR

This paper establishes a sharp early stopping rule for nonparametric testing using functional gradient descent in reproducing kernel Hilbert spaces, showing that testing optimality is achieved if and only if iterations reach an optimal order. It introduces a Wald-type test statistic based on iterated estimators and proves minimax optimality under both polynomial and exponential kernel classes.

ABSTRACT

Early stopping of iterative algorithms is an algorithmic regularization method to avoid over-fitting in estimation and classification. In this paper, we show that early stopping can also be applied to obtain the minimax optimal testing in a general non-parametric setup. Specifically, a Wald-type test statistic is obtained based on an iterated estimate produced by functional gradient descent algorithms in a reproducing kernel Hilbert space. A notable contribution is to establish a "sharp" stopping rule: when the number of iterations achieves an optimal order, testing optimality is achievable; otherwise, testing optimality becomes impossible. As a by-product, a similar sharpness result is also derived for minimax optimal estimation under early stopping studied in [11] and [19]. All obtained results hold for various kernel classes, including Sobolev smoothness classes and Gaussian kernel classes.

Motivation & Objective

  • To develop a theoretically grounded early stopping rule for nonparametric testing that avoids overfitting while maintaining minimax optimality.
  • To address the lack of a theoretically justified tuning procedure for optimal testing in nonparametric models, especially when cross-validation is suboptimal.
  • To characterize the trade-off between bias and variance in testing performance, distinct from classical estimation bias-variance trade-offs.
  • To establish a 'sharp' stopping rule where testing optimality is achieved precisely when the number of iterations reaches an optimal order.
  • To extend the sharpness result to minimax optimal estimation, showing that existing early stopping rules in [11] and [19] are also optimal for estimation.

Proposed method

  • Uses functional gradient descent in a reproducing kernel Hilbert space (RKHS) to iteratively estimate the regression function.
  • Constructs a Wald-type test statistic $ D_{n,t} $ based on the $ t $-th iterate $ f_t $, enabling hypothesis testing on the nonparametric signal.
  • Derives the strength of the weakest detectable signal (SWDS) as a function of iteration $ t $, balancing bias reduction and variance increase.
  • Introduces a data-dependent early stopping rule based on minimizing the SWDS, ensuring optimal testing performance.
  • Employs eigenvalue decay rates (polynomial and exponential) to characterize the optimal iteration count for different kernel classes.
  • Uses bootstrap approximation for bias estimation and applies theoretical bounds on shrinkage matrices and eigenvalue perturbations to derive asymptotic results.

Experimental results

Research questions

  • RQ1Can early stopping be used to achieve minimax optimal testing in nonparametric models without penalization?
  • RQ2What is the optimal number of iterations $ t $ that maximizes testing power while avoiding overfitting?
  • RQ3How does the trade-off between bias and variance in testing differ from that in estimation?
  • RQ4Is the early stopping rule for testing 'sharp'—i.e., does optimality hold only at a specific optimal iteration order?
  • RQ5Can the same stopping rule yield minimax optimal estimation, and is it sharp for estimation as well?

Key findings

  • A sharp early stopping rule is established: testing optimality is achieved if and only if the number of iterations reaches an optimal order determined by the kernel's eigenvalue decay.
  • The proposed Wald-type test statistic $ D_{n,t} $ exhibits a parabolic power pattern over iterations, peaking at the optimal $ t $, confirming the existence of a trade-off between bias and variance in testing.
  • For Sobolev smoothness classes with polynomial eigen-decay $ u_i \asymp i^{-2m} $, the optimal iteration count scales as $ t^* \asymp n^{1/(2m)} $, ensuring minimax optimality.
  • For Gaussian kernel classes with exponential eigen-decay $ u_i \asymp \exp(-\beta i^p) $, the optimal iteration count is $ t^* \asymp \log n $, again achieving minimax testing optimality.
  • The sharpness of the stopping rule is proven for both testing and estimation: optimality is unachievable if $ t \ll t^* $ or $ t \gg t^* $, with no intermediate regime offering better performance.
  • The early stopping rule derived in [11] and [19] for estimation is also shown to be sharp for minimax optimal estimation, extending their results to a broader theoretical framework.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.