[Paper Review] High-dimensional CLT for Sums of Non-degenerate Random Vectors: $n^{-1/2}$-rate
This paper establishes a high-dimensional central limit theorem (CLT) for sums of independent, non-identically distributed random vectors with non-singular covariance matrices, achieving an optimal $n^{-1/2}$ rate in the Berry–Esseen bound for hyper-rectangles. The proof leverages a non-trivial adaptation of Senatov's method of compositions, extending prior results by relaxing assumptions on sub-Gaussianity and independence while maintaining finite third moments.
In this note, we provide a Berry--Esseen bounds for rectangles in high-dimensions when the random vectors have non-singular covariance matrices. Under this assumption of non-singularity, we prove an $n^{-1/2}$ scaling for the Berry--Esseen bound for sums of mean independent random vectors with a finite third moment. The proof is essentially the method of compositions proof of multivariate Berry--Esseen bound from Senatov (2011). Similar to other existing works (Kuchibhotla et al. 2018, Fang and Koike 2020a), this note considers the applicability and effectiveness of classical CLT proof techniques for the high-dimensional case.
Motivation & Objective
- To derive a high-dimensional CLT with an $n^{-1/2}$ convergence rate for sums of independent, non-identically distributed random vectors.
- To relax restrictive assumptions such as sub-Gaussianity and log-concavity used in prior works.
- To establish the tightest possible convergence rate under minimal conditions: non-degenerate covariance and finite third moments.
- To demonstrate the optimality of the $n^{-1/2}$ rate via lower bounds, showing it cannot be improved without stronger assumptions.
- To provide a general framework applicable to high-dimensional statistical models involving multiplier random vectors, such as in regression with heavy-tailed errors.
Proposed method
- Adapts Senatov’s method of compositions to derive a multivariate Berry–Esseen bound for hyper-rectangles in high dimensions.
- Uses the Stein’s method framework to bound the difference between the distribution of the sum and a Gaussian target over rectangular sets.
- Employs a third-order Taylor expansion with controlled remainder terms via the $L^1$-norm of the third derivative of test functions.
- Introduces a normalized metric $\zeta_3(U,V)$ to quantify the discrepancy between distributions using third-order derivatives.
- Establishes a key inequality $\zeta_3(U,V) \leq \nu_3(U,V)/6$ to relate the derivative norm to the total variation-like distance.
- Constructs explicit counterexamples with heavy-tailed components to show the $n^{-1/2}$ rate is unimprovable under the given assumptions.
Experimental results
Research questions
- RQ1Can the $n^{-1/2}$ convergence rate in high-dimensional CLT be achieved under weaker assumptions than sub-Gaussianity or log-concavity?
- RQ2Is the $n^{-1/2}$ rate optimal for sums of non-i.i.d. random vectors with non-singular covariance matrices and finite third moments?
- RQ3Can the method of compositions be adapted to yield sharp bounds for hyper-rectangles in high dimensions without i.i.d. or sub-Gaussian assumptions?
- RQ4What is the minimal dependence on dimension $p$ required to achieve $n^{-1/2}$ convergence under non-degeneracy?
- RQ5How do multiplier random vectors in high-dimensional regression affect the CLT rate, and can the theory accommodate them?
Key findings
- The paper achieves an $n^{-1/2}$ Berry–Esseen bound for the sum of independent, non-identically distributed random vectors with non-singular covariance matrices and finite third moments.
- The bound holds uniformly over all $p$-dimensional hyper-rectangles, extending the applicability to high-dimensional inference.
- The $n^{-1/2}$ rate is shown to be optimal via a lower bound construction where the error remains bounded away from zero when $\sigma_{\min} \lesssim n^{-1/6}$.
- The result improves upon Lopes (2020) by removing the i.i.d. and sub-Gaussian assumptions while maintaining the $n^{-1/2}$ rate.
- The method is robust to heavy-tailed components, making it suitable for multiplier processes in high-dimensional regression with non-sub-Gaussian errors.
- The proof establishes that $\nu_3(U,V) \lesssim (\log(ep))^{3/2}$ under mild conditions, which is sufficient to achieve the $n^{-1/2}$ rate when combined with the $\sigma_{\min}$ scaling.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.