Skip to main content
QUICK REVIEW

[论文解读] A Unified Framework for Testing High Dimensional Parameters: A Data-Adaptive Approach

Cheng Zhou, Xinsheng Zhang|arXiv (Cornell University)|Aug 8, 2018
Statistical Methods and Inference参考文献 13被引用 6
一句话总结

本文提出了一种数据自适应的统一框架,用于使用 $U$-统计量和 $p \in \{1,\dots,\infty\/}$ 范围内的新型 $L_p$-范数变体进行高维假设检验,能够在多种备择情景下实现同步的检验效能。该方法引入了一种计算高效的乘子自展法,并在模拟数据和 fMRI 数据上展示了其最优性与强大的经验性能,其中数据自适应检验在注意缺陷多动障碍(ADHD)与对照组比较中获得了 $p$-值为 0.000 的结果。

ABSTRACT

High dimensional hypothesis test deals with models in which the number of parameters is significantly larger than the sample size. Existing literature develops a variety of individual tests. Some of them are sensitive to the dense and small disturbance, and others are sensitive to the sparse and large disturbance. Hence, the powers of these tests depend on the assumption of the alternative scenario. This paper provides a unified framework for developing new tests which are adaptive to a large variety of alternative scenarios in high dimensions. In particular, our framework includes arbitrary hypotheses which can be tested using high dimensional $U$-statistic based vectors. Under this framework, we first develop a broad family of tests based on a novel variant of the $L_p$-norm with $p\in \{1,\dots,\infty\}$. We then combine these tests to construct a data-adaptive test that is simultaneously powerful under various alternative scenarios. To obtain the asymptotic distributions of these tests, we utilize the multiplier bootstrap for $U$-statistics. In addition, we consider the computational aspect of the bootstrap method and propose a novel low-cost scheme. We prove the optimality of the proposed tests. Thorough numerical results on simulated and real datasets are provided to support our theory.

研究动机与目标

  • 为解决在备择假设结构未知时,缺乏一种统一且强大的高维参数检验方法的问题。
  • 开发一种能够适应高维设定下密集与稀疏备择假设的框架。
  • 确保基于自展法的 $U$-统计量在高维情形下的计算可行性与理论有效性。
  • 提供一种单一检验方法,可在不事先知晓扰动模式的情况下,保持对多样化备择情景的高检验效能。
  • 在真实 fMRI 数据上对方法进行实证验证,证明其能够检测出脑活动的组间差异。

提出的方法

  • 该框架使用 $U$-统计量来估计高维参数,对 $q$ 个参数中的每一个均采用固定阶数 $m$ 的核函数 $\Phi_s$。
  • 提出了一种针对 $p \in \{1,\dots,\infty\}$ 的新型 $L_p$-范数变体,用于构建对不同备择结构敏感的广泛检验统计量族。
  • 通过使用一种加权方案将各个 $L_p$-范数检验组合起来,形成数据自适应检验,以在不同情景下实现灵敏度的最优平衡。
  • 采用乘子自展法来近似检验统计量的渐近分布,确保在高维渐近条件下实现有效的推断。
  • 提出一种低成本的计算方案,以减轻自展法的计算负担,使其可扩展至高维数据。
  • 在较弱的矩条件与依赖性条件下证明了理论最优性,并建立了自展近似的相合性。

实验结果

研究问题

  • RQ1是否能够通过单一检验在高维参数检验中同时对密集与稀疏备择假设保持高检验效能?
  • RQ2如何构建 $L_p$-范数检验的数据自适应组合,以在多样化备择情景下维持检验效能?
  • RQ3在高维情形下,近似基于 $U$-统计量的检验统计量抽样分布的计算高效方法是什么?
  • RQ4如何对 $U$-统计量在高维设定下的乘子自展法进行适应与优化?
  • RQ5所提出的框架是否在真实世界高维数据(如 fMRI 扫描)上优于现有方法?

主要发现

  • 在检测 ADHD 与对照组之间局部低频振幅(ALFF)差异时,数据自适应检验获得了 $p$-值为 0.000,表明组间差异存在强有力的统计证据。
  • 当 $s_0 = 40$ 时,所有 $p \in \{1,\dots,\infty\}$ 的 $L_p$-范数检验均获得 $p$-值为 0.001,表明对组间差异的检测具有一致性。
  • 在对照组比较中,该方法保持了稳健性,当 $s_0 = 8000$ 时,$p$-值范围为 0.282 至 0.406,证实了在原假设下检验的有效性。
  • 所提出的低成本自展方案显著降低了计算成本,同时保持了分布近似的准确性。
  • 理论最优性已得到证明,表明该检验在各类备择模型下可达到最小最大检测率。
  • 在模拟数据与真实 fMRI 数据上的实证结果均表明,数据自适应检验在多样化情景下优于单一的 $L_p$-范数检验。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。