[论文解读] Slightly Conservative Bootstrap for Maxima of Sums
本文提出了一种略微保守的自助法,通过将自助分位数增加1%来实现高维推断中的更快收敛速度。通过使用较小的膨胀因子,覆盖误差率可提升至 $\sqrt{\log p / n}$,将样本量要求降低至 $n \gg \log p$,同时实现近乎参数化的收敛速度,显著优于标准自助法。
We study the bootstrap for the maxima of the sums of independent random variables, a problem of high relevance to many applications in modern statistics. Since the consistency of bootstrap was justified by Gaussian approximation in Chernozhukov et al. (2013), quite a few attempts have been made to sharpen the error bound for bootstrap and reduce the sample size requirement for bootstrap consistency. In this paper, we show that the sample size requirement can be dramatically improved when we make the inference slightly conservative, that is, to inflate the bootstrap quantile $t_α^*$ by a small fraction, e.g. by $1\%$ to $1.01\,t^*_α$. This simple procedure yields error bounds for the coverage probability of conservative bootstrap at as fast a rate as $\sqrt{(\log p)/n}$ under suitable conditions, so that not only the sample size requirement can be reduced to $\log p \ll n$ but also the overall convergence rate is nearly parametric. Furthermore, we improve the error bound for the coverage probability of the standard non-conservative bootstrap to $[(\log (np))^3 (\log p)^2/n]^{1/4}$ under general assumptions on data. These results are established for the empirical bootstrap and the multiplier bootstrap with third moment match. An improved coherent Lindeberg interpolation method, originally proposed in Deng and Zhang (2017), is developed to derive sharper comparison bounds, especially for the maxima.
研究动机与目标
- 为解决高维推断中 $ p \gg n $ 的挑战,特别是归一化和的最大值分布问题。
- 降低高维设定下自助法一致性的样本量要求,现有方法需满足 $ n \gg (\log p)^5 $。
- 在现有界限之外改进自助法覆盖概率误差的收敛速度,趋近于参数化速率 $ n^{-1/2} $。
- 在一般矩和尾部条件下,为经验自助法和乘子自助法提供理论保证,无需假设次高斯性。
提出的方法
- 通过将自助分位数 $ t_{\alpha}^{*} $ 增加1%至 $ 1.01\,t_{\alpha}^{*} $,提出一种略微保守的自助法,以确保更好的覆盖概率。
- 采用一致的林德贝格插值方法,推导真实分布与自助分布之间的更紧比较界,尤其适用于最大值情形。
- 在乘子自助法中应用高斯近似和三阶矩匹配,以改进误差界。
- 推导单侧覆盖误差 $ \eta^{*}_{n,\alpha} $ 的上界,表明在弱条件下其衰减速率为 $ \sqrt{\log p / n} $。
- 通过模拟研究验证在各种依赖结构和分布下,保守与精确自助法程序的性能。
- 采用经验自助法和乘子自助法方案,包括马梅恩乘子自助法和拉德马赫乘子自助法,以比较有限样本性能。
实验结果
研究问题
- RQ1在高维设定下,自助分位数的微小膨胀是否能显著改善自助法覆盖概率的收敛速度?
- RQ2当 $ p \to \infty $ 时,实现一致推断所需的最小样本量 $ n \gg \log p $ 是多少?
- RQ3在覆盖误差和收敛速度方面,保守自助法与标准(非保守)自助法相比表现如何?
- RQ4在一般矩和尾部条件下,是否可以改进标准自助法的误差界,而无需假设次高斯性?
- RQ5保守自助法在不同依赖结构和分布假设下是否保持良好的有限样本性能?
主要发现
- 略微保守自助法的覆盖误差被控制在 $ \sqrt{\log p / n} $ 以内,在一般矩和尾部条件下实现了近乎参数化的收敛速度。
- 样本量要求降低至 $ n \gg \log p $,相比先前要求 $ n \gg (\log p)^5 $ 的结果有显著改进。
- 标准(非保守)自助法的误差界被改进为 $ \big{(}(\log(np))^{3}(\log p)^{2}/n\big{)}^{1/4} $,优于早期结果。
- 采用1%膨胀的保守自助法使模拟覆盖概率平均提高约1%,在高相关性下提升更为显著。
- 马梅恩乘子自助法表现出最稳定的性能,膨胀后的分位数有效将覆盖概率推向名义水平。
- 模拟结果证实,即使标准自助法出现覆盖不足,保守自助法在相关设定下仍能实现可靠的覆盖(超过95%)。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。