Skip to main content
QUICK REVIEW

[論文レビュー] Phase Transition and Regularized Bootstrap in Large Scale $t$-tests with False Discovery Rate Control

Weidong Liu, Qi-Man Shao|arXiv (Cornell University)|Oct 16, 2013
Gene expression and cancer classification参考文献 12被引用数 9
ひとこと要約

本稿は、大規模なt検定においてp値の正規分布またはt分布近似を用いた場合の誤発見率(FDR)制御の妥当性を調査する。log m ≈ c₀n¹ᐟ³における段階的転移において、これらの近似は失敗することが判明し、有限の6次のモーメントを持つより重い尾を持つ分布に対してもFDR制御を維持できる正則化ブートストラップ法を提案する。シミュレーションでは、標準ブートストラップよりも優れた性能を示す。

ABSTRACT

Applying Benjamini and Hochberg (B-H) method to multiple Student's $t$ tests is a popular technique in gene selection in microarray data analysis. Because of the non-normality of the population, the true p-values of the hypothesis tests are typically unknown. Hence, it is common to use the standard normal distribution N(0,1), Student's $t$ distribution $t_{n-1}$ or the bootstrap method to estimate the p-values. In this paper, we first study N(0,1) and $t_{n-1}$ calibrations. We prove that, when the population has the finite 4-th moment and the dimension $m$ and the sample size $n$ satisfy $\log m=o(n^{1/3})$, B-H method controls the false discovery rate (FDR) at a given level $α$ asymptotically with p-values estimated from N(0,1) or $t_{n-1}$ distribution. However, a phase transition phenomenon occurs when $\log m\geq c_{0}n^{1/3}$. In this case, the FDR of B-H method may be larger than $α$ or even tends to one. In contrast, the bootstrap calibration is accurate for $\log m=o(n^{1/2})$ as long as the underlying distribution has the sub-Gaussian tails. However, such light tailed condition can not be weakened in general. The simulation study shows that for the heavy tailed distributions, the bootstrap calibration is very conservative. In order to solve this problem, a regularized bootstrap correction is proposed and is shown to be robust to the tails of the distributions. The simulation study shows that the regularized bootstrap method performs better than the usual bootstrap method.

研究の動機と目的

  • 大規模な多重仮説検定における標準正規分布またはt分布近似を用いたFDR制御の漸近的挙動を分析すること。
  • 高次元設定における段階的転移のための、これらの近似が失敗する条件を同定すること。
  • 重い尾を持つ分布においてもFDR制御を維持できる、ブートストラップ補正の代替手法を開発すること。
  • 有限6次モーメントのみを要件とする正則化ブートストラップ法を提案し、理論的に正当化すること。標準ブートストラップの過剰な保守性を改善する。

提案手法

  • 大規模t検定におけるp値の正規分布およびt分布近似の下でのFDR制御の理論的分析。
  • 次元mと標本サイズnの関係に基づく段階的転移閾値の導出、特にlog m ≥ c₀n¹ᐟ³のとき。
  • 特にサブガウス尾を持つ場合にp値推定の精度を向上させるためのブートストラップ補正の使用。
  • 極端な観測値を切断することで重い尾への感受性を低減する正則化ブートストラップ法の導入。
  • 集中不等式とモーメントバウンドを用いて、切断の下での経験モーメントの収束を証明。
  • 正則化ブートストラップにおけるFDR制御の理論的正当化。log m = o(n¹ᐟ²)の下で漸近的に有効であることを示す。

実験結果

リサーチクエスチョン

  • RQ1標準正規分布およびt分布近似を用いたp値のFDR制御が、大規模t検定でいつ失敗するか。
  • RQ2log m ≥ c₀n¹ᐟ³のとき、FDR制御における段階的転移の性質と閾値は何か。
  • RQ3ブートストラップ補正の性能は尾の挙動にどのように依存するか。重い尾を持つ分布に対しても改善可能か。
  • RQ4有限6次モーメントを持つより重い尾を持つ分布に対しても、正則化ブートストラップ法がFDR制御を維持できるか。
  • RQ5有限サンプルのシミュレーションにおいて、重い尾を持つデータに対して、正則化ブートストラップ法は標準ブートストラップよりもロバストか。

主な発見

  • log m = o(n¹ᐟ³)のとき、4次のモーメントが有限であれば、標準正規分布またはt_{n-1}のp値近似を用いたBenjamini-Hochberg法によるFDR制御は漸近的に有効である。
  • log m ≥ c₀n¹ᐟ³のとき段階的転移が発生する。非ゼロの歪度がある場合、FDRはαを超える可能性があり、log m/n¹ᐟ³ → ∞ に伴い1に近づくことがある。
  • ブートストラップ補正は、log m = o(n¹ᐟ²)の下で、元の分布がサブガウス尾を持つ限り、FDR制御を維持する。
  • 標準ブートストラップは、モーメントが有限であっても、重い尾を持つ分布に対して過剰に保守的になる。
  • 提案された正則化ブートストラップ法は、log m = o(n¹ᐟ²)の下で有限6次モーメントのみを要件とし、FDR制御を維持する。シミュレーションでは標準ブートストラップを上回る性能を示す。
  • 理論的およびシミュレーション結果により、正則化ブートストラップ法は尾の重さに対してロバストであり、重い尾の設定において標準ブートストラップよりもより正確なFDR制御を提供することが確認された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。