Skip to main content
QUICK REVIEW

[论文解读] To how many simultaneous hypothesis tests can normal, Student's t or bootstrap calibration be applied?

Jianqing Fan, P. Hall|ArXiv.org|Dec 29, 2006
Statistical Methods in Clinical Trials参考文献 18被引用 11
一句话总结

本文確立了使用常態、學生t或自助法校準時,可準確進行同時假設檢定之數量的理論極限。結果顯示,使用常態或t校準時,log N 必須比 n^{1/3} 增長得更慢;而使用自助法校準時,log N 可幾乎與 n^{1/2} 同樣快速增長,於微陣列分析等高維設定中具有顯著優勢。

ABSTRACT

In the analysis of microarray data, and in some other contemporary statistical problems, it is not uncommon to apply hypothesis tests in a highly simultaneous way. The number, $ν$ say, of tests used can be much larger than the sample sizes, $n$, to which the tests are applied, yet we wish to calibrate the tests so that the overall level of the simultaneous test is accurate. Often the sampling distribution is quite different for each test, so there may not be an opportunity for combining data across samples. In this setting, how large can $ν$ be, as a function of $n$, before level accuracy becomes poor? In the present paper we answer this question in cases where the statistic under test is of Student's $t$ type. We show that if either Normal or Student's $t$ distribution is used for calibration then the level of the simultaneous test is accurate provided $\logν$ increases at a strictly slower rate than $n^{1/3}$ as $n$ diverges. On the other hand, if bootstrap methods are used for calibration then we may choose $\logν$ almost as large as $n\half$ and still achieve asymptotic level accuracy. The implications of these results are explored both theoretically and numerically.

研究动机与目标

  • 確定當 N ≫ n(樣本大小)時,可準確校準的同時假設檢定數量 N 的最大值。
  • 評估在高維設定下,常態、學生t與自助法校準方法的漸近顯著水準準確性。
  • 建立 log N 對 n 的理論邊界,以維持準確的家族錯誤率或偽發現在率控制。
  • 比較不同校準方法在 N 增加時的可擴展性表現。

提出的方法

  • 作者使用埃傑沃斯展開與大偏差展開,分析常態、t與自助法近似下檢定統計量的尾部行為。
  • 透過經驗矩量推導 t 統計量及其自助對應物的漸近展開,並以經驗累積量引入偏態與峰態項。
  • 該方法依賴伯恩斯坦不等式與經驗過程的均勻收斂性,以控制樣本矩與母體矩之間的偏離。
  • 透過分析不同校準方案下檢定統計量分位數估計的誤差,推導理論邊界。
  • 分析涵蓋抽樣分配的中央與尾部行為,專注於同時推論中臨界值的準確性。
  • 證明顯示,當 log N 增長為 o(n^{1/2}) 時,自助法校準能維持準確性,而常態/t校準則需 log N = o(n^{1/3})。

实验结果

研究问题

  • RQ1當 N ≫ n 時,可使用可接受的顯著水準準確性校準的同時假設檢定數量 N 的最大值為何?
  • RQ2校準方法的選擇——常態、t 或自助法——如何影響高維設定下同時推論的可擴展性?
  • RQ3在抽樣分配具有偏態或重尾等條件下,常態或 t 校準對大 N 仍有效的條件為何?
  • RQ4當 log N 几乎與 n^{1/2} 同樣快速增長時,自助法校準能否維持漸近顯著水準準確性?
  • RQ5P值估計所需的準確性如何隨 N 變化?這對偽發現在率控制有何影響?

主要发现

  • 使用常態或學生t校準時,即使在偏態或重尾分配下,漸近顯著水準準確性僅在 log N = o(n^{1/3}) 時能維持。
  • 自助法校準允許在 log N = o(n^{1/2}) 時維持漸近顯著水準準確性,與常態/t校準相比顯著放寬了限制。
  • 自助法校準的改善來自其更能捕捉抽樣分配的偏態與高階矩的特性。
  • 這些結果在最小的矩條件下仍成立:常態/t校準僅需有界第三矩,而自助法校準則需更強的矩限制。
  • 理論邊界獲得數值探勘的支持,顯示當 N 接近 n^{1/2} 時,自助法校準仍保持準確。
  • 研究結果表示,與漸近近似相比,自助法在基因體學等領域的大規模多重檢定中遠為適用。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。