Skip to main content
QUICK REVIEW

[论文解读] Incorporation of Sparsity Information in Large-scale Multiple Two-sample $t$ Tests

Weidong Liu|arXiv (Cornell University)|Oct 16, 2014
Statistical Methods in Clinical Trials参考文献 22被引用 7
一句话总结

本文提出了一种基于非相关筛选(US)的程序,用于大规模多重两样本t检验,通过利用均值向量中的稀疏性来提高统计功效,同时控制错误发现率(FDR)。通过使用原始数据构建渐近上与t统计量不相关的筛选统计量——无需样本分割——该方法在稀疏性假设下,其功效高于Benjamini-Hochberg方法,且渐近上保证了FDR控制。

ABSTRACT

Large-scale multiple two-sample {\em Student}'s $t$ testing problems often arise from the statistical analysis of scientific data. To detect components with different values between two mean vectors, a well-known procedure is to apply the Benjamini and Hochberg (B-H) method and two-sample {\em Student}'s $t$ statistics to control the false discovery rate (FDR). In many applications, mean vectors are expected to be sparse or asymptotically sparse. When dealing with such type of data, {\em can we gain more power than the standard procedure such as the B-H method with Student's $t$ statistics while keeping the FDR under control?} The answer is positive. By exploiting the possible sparsity information in mean vectors, we present an uncorrelated screening-based (US) FDR control procedure, which is shown to be more powerful than the B-H method. The US testing procedure depends on a novel construction of screening statistics, which are asymptotically uncorrelated with two-sample {\em Student}'s $t$ statistics. The US testing procedure is different from some existing {\em testing following screening} methods (Reiner, et al., 2007; Yekutieli, 2008) in which independence between screening and testing is crucial to control the FDR, while the independence often requires additional data or splitting of samples. An inappropriate splitting of samples may result in a loss rather than an improvement of statistical power. Instead, the uncorrelated screening US is based on the original data and does not need to split the samples. Theoretical results show that the US testing procedure controls the desired FDR asymptotically. Numerical studies are conducted and indicate that the proposed procedure works quite well.

研究动机与目标

  • 解决当均值向量稀疏或渐近稀疏时,大规模多重两样本t检验中统计功效较低的挑战。
  • 开发一种在稀疏性下优于标准Benjamini-Hochberg(B-H)程序的方法,同时保持FDR控制。
  • 通过构建在分布上与检验统计量独立的筛选统计量,避免对样本分割或额外数据的需求。
  • 在高维稀疏均值差异设定下,建立FDR控制和功效提升的理论保证。

提出的方法

  • 提出一种新颖的非相关筛选(US)统计量,其渐近上与两样本t统计量不相关,从而可在不违反FDR假设的前提下联合使用。
  • 使用原始数据同时进行筛选和检验,消除了样本分割的需求,避免了因数据划分导致的功效损失。
  • 基于筛选统计量设计阈值规则,识别候选假设进行检验,从而在保持FDR控制的同时减少检验数量。
  • 对筛选后的假设应用修改后的Benjamini-Hochberg程序,利用筛选集中t检验的p值的顺序统计量。
  • 在稀疏性假设下推导筛选统计量和检验统计量的渐近分布,确保FDR在名义水平α上保持有界。
  • 建立理论条件,表明在真实信号大小较小时且稀疏时,US程序的功效高于标准B-H程序。

实验结果

研究问题

  • RQ1当均值向量稀疏时,能否在不损害FDR控制的前提下,提升大规模多重两样本t检验的统计功效?
  • RQ2是否可以构建一个渐近上与t统计量不相关的筛选统计量,从而在无需样本分割的情况下实现FDR控制中的联合使用?
  • RQ3所提出的非相关筛选(US)程序在何种理论条件下可实现FDR控制并优于B-H方法?
  • RQ4US程序的性能与现有依赖样本分割或独立性假设的筛选后检验方法相比如何?

主要发现

  • 所提出的非相关筛选(US)程序在高维稀疏性下,渐近地在名义水平α上控制了错误发现率(FDR)。
  • 当均值向量稀疏或渐近稀疏时,US方法的统计功效高于标准Benjamini-Hochberg(B-H)程序。
  • 筛选统计量被构造为渐近上与t统计量不相关,从而可直接使用原始数据而无需分割,避免了因数据划分导致的功效损失。
  • 数值研究证实,US程序在有限样本下表现良好,能够在稀疏性下实现FDR控制并提升检测功效。
  • 理论分析表明,只要信号强度以适当速率增长,US程序的功效在概率上随维度m增加而趋近于1。
  • 该方法对未知的稀疏模式具有鲁棒性,且无需事先知晓非零均值差异的联合支持集。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。