Skip to main content
QUICK REVIEW

[Paper Review] Notes on the dimension dependence in high-dimensional central limit theorems for hyperrectangles

Yuta Koike|arXiv (Cornell University)|Nov 1, 2019
Stochastic processes and statistical mechanics44 references22 citations
TL;DR

This paper improves high-dimensional central limit theorems for hyperrectangles by reducing the dimension dependence in approximation error bounds. Under moment conditions, it shows that the probability that a normalized sum of i.i.d. random vectors lies in a hyperrectangle can be uniformly approximated by its Gaussian counterpart as n → ∞, with the error decaying when (log d)^5 / n → 0. When a common factor structure is present, the condition improves to (log d)^3 / n → 0.

ABSTRACT

Let $X_1,\dots,X_n$ be independent centered random vectors in $\mathbb{R}^d$. This paper shows that, even when $d$ may grow with $n$, the probability $P(n^{-1/2}\sum_{i=1}^nX_i\in A)$ can be approximated by its Gaussian analog uniformly in hyperrectangles $A$ in $\mathbb{R}^d$ as $n o\infty$ under appropriate moment assumptions, as long as $(\log d)^5/n o0$. This improves a result of Chernozhukov, Chetverikov & Kato [Ann. Probab. 45 (2017) 2309-2353] in terms of the dimension growth condition. When $n^{-1/2}\sum_{i=1}^nX_i$ has a common factor across the components, this condition can be further improved to $(\log d)^3/n o0$. The corresponding bootstrap approximation results are also developed. These results serve as a theoretical foundation of simultaneous inference for high-dimensional models.

Motivation & Objective

  • To improve the dimension growth condition required for uniform Gaussian approximation of high-dimensional sums over hyperrectangles.
  • To address the open question of whether the log^7 d dependence in Chernozhukov et al. (2017) can be reduced.
  • To establish tighter bounds for simultaneous inference in high-dimensional models.
  • To develop bootstrap approximation results that complement the main CLT bounds.

Proposed method

  • Uses a randomized Lindeberg method with Stein kernel techniques to control the dependence of the approximation error on dimension d.
  • Introduces a concentration function CF(ε) and its dual measure ΘX to quantify anti-concentration of the Gaussian limit.
  • Applies anti-concentration inequalities and moment bounds on fourth-order moments and ψ1-norms of components.
  • Derives bounds on the suprema of differences between P(SX_n ∈ A) and P(ZX_n ∈ A) over all hyperrectangles A.
  • Develops a bootstrap approximation theorem by coupling the main CLT result with a symmetrization and conditioning argument.
  • Employs a symmetrized version of the sum (X⋄) to decouple the dependence on the empirical mean and control the error via concentration inequalities.

Experimental results

Research questions

  • RQ1Can the dimension growth condition (log d)^7 / n → 0 in Chernozhukov et al. (2017) be improved for hyperrectangle approximation?
  • RQ2Does the presence of a common factor structure in the normalized sum allow for a tighter dimension dependence?
  • RQ3Is the (log d)^5 / n condition optimal in the general case, or can it be further reduced?
  • RQ4Can the bootstrap approximation for maxima over components be consistently applied under the same improved dimension condition?
  • RQ5What is the role of anti-concentration (measured by ΘX) in enabling high-dimensional CLTs for non-convex sets like hyperrectangles?

Key findings

  • The paper establishes that ρn(Are) → 0 as n → ∞ under the condition (log d)^5 / n → 0, improving upon the (log d)^7 / n → 0 condition in Chernozhukov et al. (2017).
  • When the normalized sum SX_n has a common factor across components, the condition improves to (log d)^3 / n → 0.
  • The bound is of the form ρn(Are) ≲ Θ_X^{2/3} (B_n^4 log^3 d / n)^{1/6}, where Θ_X measures anti-concentration and B_n controls moment growth.
  • The bootstrap approximation error is shown to be uniformly bounded under the same (log d)^5 / n → 0 condition, enabling valid inference in high-dimensional models.
  • The improvement is achieved via a refined application of the randomized Lindeberg method and sharper anti-concentration control using the Stein kernel and ψ1-norms.
  • The results provide a theoretical foundation for simultaneous inference, such as confidence intervals and multiple testing with strong family-wise error rate control.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.