[论文解读] Dimension-agnostic inference using cross U-statistics
该论文提出了一种基于交叉U-统计量的维度无关推理框架,无论维度 $d$ 如何随样本量 $n$ 变化,该框架均能实现高斯极限分布。通过利用样本分割、自标准化以及检验统计量的变分表示,该方法在固定、高维及超高维情形下均能实现有效的推断,其检验功效与针对特定情形的U-统计量相比,仅相差 $\sqrt{2}$ 因子。
Classical asymptotic theory for statistical inference usually involves calibrating a statistic by fixing the dimension $d$ while letting the sample size $n$ increase to infinity. Recently, much effort has been dedicated towards understanding how these methods behave in high-dimensional settings, where $d$ and $n$ both increase to infinity together. This often leads to different inference procedures, depending on the assumptions about the dimensionality, leaving the practitioner in a bind: given a dataset with 100 samples in 20 dimensions, should they calibrate by assuming $n \gg d$, or $d/n \approx 0.2$? This paper considers the goal of dimension-agnostic inference; developing methods whose validity does not depend on any assumption on $d$ versus $n$. We introduce an approach that uses variational representations of existing test statistics along with sample splitting and self-normalization to produce a refined test statistic with a Gaussian limiting distribution, regardless of how $d$ scales with $n$. The resulting statistic can be viewed as a careful modification of degenerate U-statistics, dropping diagonal blocks and retaining off-diagonal blocks. We exemplify our technique for some classical problems including one-sample mean and covariance testing, and show that our tests have minimax rate-optimal power against appropriate local alternatives. In most settings, our cross U-statistic matches the high-dimensional power of the corresponding (degenerate) U-statistic up to a $\sqrt{2}$ factor.
研究动机与目标
- 解决统计推断中的实际困境:实践中研究者必须根据 $d$ 相对于 $n$ 的假设选择校准方法。
- 开发一种单一的推断程序,无论维度 $d$ 如何随样本量 $n$ 变化,其有效性均不受影响。
- 在不同渐近情形下实现极小极大率最优功效,且无需事先知晓 $d$ 与 $n$ 的缩放关系。
- 提出一种可推广的方法论框架,用于维度无关的推断,结合变分表示与样本分割。
提出的方法
- 使用样本分割将数据划分为互不相交的子集,以实现条件期望的独立估计。
- 对U-统计量应用自标准化,以稳定方差并确保其收敛至高斯分布。
- 利用检验统计量的变分表示,解耦其对维度的依赖性。
- 通过舍弃对角块并仅保留非对角U-统计量分量,构建改进的检验统计量。
- 在一般 $d_n/n$ 缩放下,推导交叉U-统计量的渐近正态性,涵盖固定、高维及超高维情形。
- 结合马尔可夫不等式与切比雪夫不等式,并辅以精细的浓度不等式,统一控制第一类与第二类错误。

实验结果
研究问题
- RQ1是否存在一种单一统计检验,可在所有 $d_n$ 相对于 $n$ 的缩放情形下(包括固定、高维及超高维情形)保持渐近有效性?
- RQ2是否可能设计出一种推断程序,其功效能自适应最优地匹配潜在维度,且无需事先知晓缩放关系?
- RQ3是否可通过精心构造,使检验统计量的极限分布在任意 $d/n$ 情形下均近似为高斯分布?
- RQ4维度无关推断的功效与针对特定情形的U-统计量方法相比如何?
- RQ5有哪些通用的方法论原则可使维度无关推断超越特定检验统计量而具有普适性?
主要发现
- 所提出的交叉U-统计量在所有 $d_n/n$ 缩放情形下均达到高斯极限分布,从而实现通用校准。
- 该方法确保了在任意维度假设下,显著性水平为 $\alpha$ 的检验与置信度为 $1-\alpha$ 的置信区间均具有渐近有效性。
- 该检验在局部替代假设下保持极小极大率最优功效,其性能与退化U-统计量相比仅相差 $\sqrt{2}$ 因子。
- 第一类错误在所有情形下均得到统一控制,且当 $\min\{m_1, m_2\} \to \infty$ 时,第二类错误趋于零。
- 该方法在无需知晓 $d_n$ 或其相对于 $n$ 的增长速率下,实现了统一的渐近有效性。
- 通过浓度不等式与变分表示建立了理论保证,确保在多种高维设定下具有鲁棒性。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。