[论文解读] High-dimensional Central Limit Theorems by Stein's Method
该论文利用 Stein 方法,针对超长方体上的正态近似,推导出高维中心极限定理的显式误差界。对于具有对数凹密度的独立同分布各向同性随机向量,其误差率最优,仅相差 log n 因子,该结果基于 Götze 的方法,结合 Anderson 等人(1998)的改进估计,以及 Bhattacharya 和 Rao(1976)的平滑不等式。
We obtain explicit error bounds for the $d$-dimensional normal approximation on hyperrectangles for a random vector that has a Stein kernel, or admits an exchangeable pair coupling, or is a non-linear statistic of independent random variables or a sum of $n$ locally dependent random vectors. We assume the approximating normal distribution has a non-singular covariance matrix. The error bounds vanish even when the dimension $d$ is much larger than the sample size $n$. We prove our main results using the approach of Götze (1991) in Stein's method, together with modifications of an estimate of Anderson, Hall and Titterington (1998) and a smoothing inequality of Bhattacharya and Rao (1976). For sums of $n$ independent and identically distributed isotropic random vectors having a log-concave density, we obtain an error bound that is optimal up to a $\log n$ factor. We also discuss an application to multiple Wiener-Itô integrals.
研究动机与目标
- 解决当维度 d 远大于样本量 n 时的高维正态近似挑战。
- 在一般依赖结构下,为超长方体上的多变量正态近似提供显式误差界。
- 改进文献中已有的误差界,特别是针对具有对数凹密度的独立同分布各向同性向量。
- 将 Stein 方法的应用范围扩展至具有 Stein 核、交换对或局部依赖结构的随机向量。
- 建立一个框架,即使在 d ≫ n 时也能实现误差界趋于零,从而支持高维统计推断。
提出的方法
- 采用 Götze(1991)在多变量正态近似中应用 Stein 方法的思路。
- 引入 Anderson、Hall 和 Titterington(1998)的改进估计,以控制依赖性和尾部行为。
- 应用 Bhattacharya 和 Rao(1976)的平滑不等式,以处理超长方体上的上确界。
- 利用 Stein 核表示法刻画随机向量 W 的分布,而无需假设独立同分布或独立分量。
- 通过交换对和局部依赖结构的耦合技术,将结果推广至非独立同分布情形。
- 应用高斯测度的反浓度不等式,以控制扰动集合上概率的差异。
实验结果
研究问题
- RQ1当 d ≫ n 时,能否为高维正态近似在超长方体上导出显式误差界?
- RQ2Stein 方法如何在高维中扩展至具有 Stein 核或交换对耦合的随机向量?
- RQ3具有对数凹密度的独立同分布各向同性向量在高维正态近似中的最优收敛速率为何?
- RQ4能否通过更精细的估计和平滑技术,将误差界改进至超过现有文献结果?
- RQ5协方差矩阵的最小特征值在决定高维中心极限定理收敛速率中起什么作用?
主要发现
- 对于协方差非奇异的一般随机向量,论文得到误差界为 $ C_0 \big( \frac{\tau^4 \text{log}^4(dn)}{n} \big)^{1/3} $,优于 Chernozhukov 等人(2019)的 $ ( \text{log}^5(dn)/n )^{1/4} $ 界。
- 对于具有对数凹密度的独立同分布各向同性随机向量,误差界为 $ C_0 \big( \frac{\text{log}^3 d}{n} \big)^{1/2} \text{log} n $,仅相差 $ \text{log} n $ 因子,为最优。
- 即使当 $ d \gg n $ 时,误差界仍趋于零,从而支持有效的高维推断。
- 该方法适用于广泛依赖结构,包括局部依赖向量和非线性统计量。
- 该方法在弱矩条件下方可适用,且不要求维度相对于样本量为次线性。
- 建立了该方法在多重 Wiener-Itô 积分中的应用,展示了其在随机分析中的实用性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。