[论文解读] Gaussian approximations and multiplier bootstrap for maxima of sums of high-dimensional random vectors
该论文为高维随机向量和的最大值发展了高斯近似和乘子自展法,即使维度 $ p $ 远大于样本量 $ n $ 时也适用。它建立了非渐近界,表明最大值的真实分布与高斯或条件高斯近似之间的柯尔莫哥洛夫距离随 $ n $ 多项式衰减,从而在高维统计中实现有效的推断。
We derive a Gaussian approximation result for the maximum of a sum of high-dimensional random vectors. Specifically, we establish conditions under which the distribution of the maximum is approximated by that of the maximum of a sum of the Gaussian random vectors with the same covariance matrices as the original vectors. This result applies when the dimension of random vectors ($p$) is large compared to the sample size ($n$); in fact, $p$ can be much larger than $n$, without restricting correlations of the coordinates of these vectors. We also show that the distribution of the maximum of a sum of the random vectors with unknown covariance matrices can be consistently estimated by the distribution of the maximum of a sum of the conditional Gaussian random vectors obtained by multiplying the original vectors with i.i.d. Gaussian multipliers. This is the Gaussian multiplier (or wild) bootstrap procedure. Here too, $p$ can be large or even much larger than $n$. These distributional approximations, either Gaussian or conditional Gaussian, yield a high-quality approximation to the distribution of the original maximum, often with approximation error decreasing polynomially in the sample size, and hence are of interest in many applications. We demonstrate how our Gaussian approximations and the multiplier bootstrap can be used for modern high-dimensional estimation, multiple hypothesis testing, and adaptive specification testing. All these results contain nonasymptotic bounds on approximation errors.
研究动机与目标
- 解决当维度 $ p $ 相对于样本量 $ n $ 较大时,高维随机向量和的最大值的分布近似问题。
- 为独立同分布的高维随机向量和的最大分量的分布提供一个非渐近高斯近似。
- 开发一种有效的乘子自展程序,用于在协方差结构未知时估计最大值的分布。
- 建立近似误差随 $ n $ 多项式衰减的条件,即使 $ p $ 随 $ n $ 指数增长也成立。
- 使高维统计问题中的实际推断成为可能,如多重检验、估计和模型设定检验。
提出的方法
- 通过将其与具有相同协方差矩阵的高斯向量和的最大值进行比较,推导出高维随机向量和的最大值的高斯近似。
- 使用 Stein 方法、Slepian 插值法和反浓度不等式,控制真实分布与近似分布之间的柯尔莫哥洛夫距离。
- 对最大函数应用光滑势近似,将其与自旋玻璃模型中的自由能联系起来,以增强分析可处理性。
- 通过将原始向量乘以独立同分布的标准正态乘子,引入高斯乘子(野蛮)自展法,以估计未知的基于协方差的分布。
- 利用矩条件和反浓度条件,建立近似误差的非渐近界,包括次指数尾部和非次指数尾部。
- 通过浓度不等式和自标准化和界,控制自展法中的误差,即限制经验协方差结构与真实协方差结构之间的差异。
实验结果
研究问题
- RQ1在何种条件下,高维随机向量和的最大值的分布会收敛到具有相同协方差结构的高斯向量分布?
- RQ2当真实协方差矩阵未知时,能否使用乘子自展法一致地估计最大值的分布?
- RQ3近似误差在样本量 $ n $ 中衰减的速度如何,特别是在 $ p \gg n $ 时?
- RQ4哪些矩条件和依赖结构条件能保证高维设置下高斯近似与自展法的有效性?
- RQ5这些近似方法如何应用于高维推断问题,如多重假设检验和自适应模型设定检验?
主要发现
- 在较弱的矩条件下,最大值 $ T_0 $ 的真实分布与其中高斯近似 $ Z_0 $ 之间的柯尔莫哥洛夫距离被控制在 $ Cn^{-c} $ 以内,其中 $ c>0 $。
- 即使 $ p $ 的增长快至 $ e^{o(n^c)} $,近似误差仍随 $ n $ 多项式衰减,从而允许 $ p \gg n $。
- 对于有界随机向量,误差衰减速率为 $ O((\log p)^7 / n) $,只要 $ (\log p)^7 / n \to 0 $,该速率即趋于零。
- 在相同条件下,高斯乘子自展法能一致地估计最大值的分布,其误差率同样为 $ O(n^{-c}) $。
- 该方法在尾部概率近似中表现出高精度,这对实际推断(如多重检验)至关重要。
- 结果为非渐近性,且在高维设置中保持一致,使现代数据分析中 $ p \gg n $ 的有效推断成为可能。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。