[论文解读] Randomized incomplete $U$-statistics in high dimensions
该论文提出了一种使用稀疏、随机加权的随机不完整U统计量,以实现高维U统计量的计算高效推断,其中样本量 $ n $ 和维度 $ d $ 均较大。该方法实现了非渐近高斯近似误差界,并提出了一种在高维下计算高效的自助法程序,即使在 $ d \gg n $ 的情况下也适用于退化和非退化核函数。
This paper studies inference for the mean vector of a high-dimensional $U$-statistic. In the era of Big Data, the dimension $d$ of the $U$-statistic and the sample size $n$ of the observations tend to be both large, and the computation of the $U$-statistic is prohibitively demanding. Data-dependent inferential procedures such as the empirical bootstrap for $U$-statistics is even more computationally expensive. To overcome such computational bottleneck, incomplete $U$-statistics obtained by sampling fewer terms of the $U$-statistic are attractive alternatives. In this paper, we introduce randomized incomplete $U$-statistics with sparse weights whose computational cost can be made independent of the order of the $U$-statistic. We derive non-asymptotic Gaussian approximation error bounds for the randomized incomplete $U$-statistics in high dimensions, namely in cases where the dimension $d$ is possibly much larger than the sample size $n$, for both non-degenerate and degenerate kernels. In addition, we propose generic bootstrap methods for the incomplete $U$-statistics that are computationally much less-demanding than existing bootstrap methods, and establish finite sample validity of the proposed bootstrap methods. Our methods are illustrated on the application to nonparametric testing for the pairwise independence of a high-dimensional random vector under weaker assumptions than those appearing in the literature.
研究动机与目标
- 解决在 $ d \gg n $ 且 $ r \geq 3 $ 的高维设置下完整U统计量的计算不可行性问题。
- 为U统计量开发一种计算高效的替代方法,以替代经验自助法,因为后者由于需要 $ O(Bn^r d) $ 次运算而计算成本过高。
- 为高维情形下的不完整U统计量建立非渐近高斯近似误差界,涵盖退化和非退化情形。
- 提出一种通用的、计算高效的不完整U统计量自助法,具有有限样本有效性。
- 展示该方法在较弱矩假设下对成对独立性进行非参数检验的实用性,优于以往工作。
提出的方法
- 通过伯努利抽样和有放回抽样构造稀疏权重,引入使用随机抽样的不完整U统计量,将计算成本与 $ r $ 解耦。
- 采用分而治之策略和随机抽样估计,高效近似自助分布。
- 推导出高维情形下随机不完整U统计量的非渐近高斯近似误差界,明确依赖于 $ n $、$ d $ 和 $ r $。
- 提出一种通用的自助法——MB-NDG-DC 和 MB-NDG-RS——基于不完整U统计量,理论证明其具有有限样本有效性。
- 采用一种针对随机抽样设计进行调整的标准化方案,提升方差控制和近似精度。
- 采用线性模型分析计算运行时间,表明自助法的运行时间呈 $ O(n^2) $ 阶,与理论预测一致。
实验结果
研究问题
- RQ1在 $ d \gg n $ 的高维设置下,随机不完整U统计量能否实现准确的高斯近似?
- RQ2在高维情形下,随机不完整U统计量的高斯近似的非渐近误差界是什么?
- RQ3能否为不完整U统计量构建一种计算高效的自助程序,同时保持有限样本有效性?
- RQ4在计算成本和准确性方面,所提自助法与标准经验自助法相比表现如何?
- RQ5该方法能否应用于在弱于现有方法的矩假设下进行成对独立性的非参数检验?
主要发现
- 随机不完整U统计量实现了非渐近高斯近似误差界,其误差界在 $ n $、$ d $ 和 $ r $ 上具有有利的依赖关系,即使在高维情形 $ d \gg n $ 下也成立。
- 所提出的自助法(MB-NDG-DC 和 MB-NDG-RS)将计算成本从 $ O(Bn^r d) $ 降低至 $ O(BN d) $,其中 $ N \ll n^r $,运行时间呈 $ O(n^2) $ 阶,与理论预测一致。
- 实证结果表明,所有测试配置下的自助检验的均匀大小误差均低于 0.02,表明其具有强大的有限样本有效性。
- 高斯近似非常精确,P-P图显示 $ \sqrt{n}U_{n,N}' $ 的经验分布与目标正态分布高度吻合。
- 随机标准化方案相比确定性标准化,使近似正态分布的方差更小,从而提升了近似精度。
- 该方法成功实现了在弱于以往要求的矩条件下对成对独立性的非参数检验,如在 $ a = 0.9 $ 的Copula相关性示例中所展示。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。