[論文レビュー] Incorporation of Sparsity Information in Large-scale Multiple Two-sample $t$ Tests
本稿では、平均ベクトルのスパarsityを活用して統計的パワーを向上させながら、誤り発見率(FDR)を制御する大規模な複数の2標本t検定に対して、相関のないスクリーニング(US)手法を提案する。元のデータを用いて、標本分割を伴わずにt統計量と漸近的に無相関となるスクリーニング統計量を構築することで、スパarsity仮定下でBenjamini-Hochberg法よりも高いパワーを達成し、FDR制御が漸近的に保証される。
Large-scale multiple two-sample {\em Student}'s $t$ testing problems often arise from the statistical analysis of scientific data. To detect components with different values between two mean vectors, a well-known procedure is to apply the Benjamini and Hochberg (B-H) method and two-sample {\em Student}'s $t$ statistics to control the false discovery rate (FDR). In many applications, mean vectors are expected to be sparse or asymptotically sparse. When dealing with such type of data, {\em can we gain more power than the standard procedure such as the B-H method with Student's $t$ statistics while keeping the FDR under control?} The answer is positive. By exploiting the possible sparsity information in mean vectors, we present an uncorrelated screening-based (US) FDR control procedure, which is shown to be more powerful than the B-H method. The US testing procedure depends on a novel construction of screening statistics, which are asymptotically uncorrelated with two-sample {\em Student}'s $t$ statistics. The US testing procedure is different from some existing {\em testing following screening} methods (Reiner, et al., 2007; Yekutieli, 2008) in which independence between screening and testing is crucial to control the FDR, while the independence often requires additional data or splitting of samples. An inappropriate splitting of samples may result in a loss rather than an improvement of statistical power. Instead, the uncorrelated screening US is based on the original data and does not need to split the samples. Theoretical results show that the US testing procedure controls the desired FDR asymptotically. Numerical studies are conducted and indicate that the proposed procedure works quite well.
研究の動機と目的
- 平均ベクトルがスパースまたは漸近的にスパースである場合に、大規模な複数の2標本t検定において統計的パワーが低いという課題に対処すること。
- スパarsity下でFDR制御を維持しつつ、標準的なBenjamini-Hochberg(B-H)手順よりもパワーを向上させる手法を開発すること。
- 分布上、スクリーニング統計量と検定統計量が独立であるように構築することで、標本分割や追加データの必要性を回避すること。
- 高次元設定におけるスパースな平均差の下で、FDR制御とパワー向上の理論的保証を確立すること。
提案手法
- 2標本t統計量と漸近的に無相関となる新しい相関のないスクリーニング(US)統計量を導入し、FDR仮定を破らない形で併用可能となる。
- スクリーニングと検定の両方に元のデータを用いることで、標本分割の必要性がなくなり、データ分割によるパワー損失を回避する。
- スクリーニング統計量に基づくしきい値ルールを用いて、検定の対象となる候補仮説を特定し、検定回数を削減しながらFDR制御を維持する。
- スクリーニングされた集合におけるt検定のp値の順序統計量を用いて、修正されたBenjamini-Hochberg手順を適用する。
- スパarsity下でスクリーニング統計量および検定統計量の漸近的分布を導出し、FDRが名目水準αで有界であることを保証する。
- 真の信号サイズが小さくスパースな場合に、US手順が標準的なB-H手順よりも高いパワーを達成する理論的条件を確立する。
実験結果
リサーチクエスチョン
- RQ1平均ベクトルがスパースである場合に、FDR制御を損なわずに大規模な複数の2標本t検定における統計的パワーを向上させることは可能か?
- RQ2t統計量と漸近的に無相関となるスクリーニング統計量を構築することは可能か?これにより、標本分割なしにFDR制御における併用が可能になるか?
- RQ3提案された相関のないスクリーニング(US)手順がFDRを制御し、B-H法よりも高いパワーを達成するための理論的条件は何か?
- RQ4US手順の性能は、標本分割や独立性仮定に依存する既存のスクリーニング後に検定を行う手法と比べてどのように異なるか?
主な発見
- 提案された相関のないスクリーニング(US)手順は、高次元的スパarsity下でも、漸近的に名目水準αで誤り発見率(FDR)を制御する。
- 平均ベクトルがスパースまたは漸近的にスパースである場合、US法は標準的なBenjamini-Hochberg(B-H)手順よりも高い統計的パワーを達成する。
- スクリーニング統計量はt統計量と漸近的に無相関となるように構築されており、元のデータを分割せずに使用可能であるため、データ分割によるパワー損失を回避できる。
- 数値的実験では、US手順が有限標本でも良好に機能し、スパarsity下でFDR制御と改善された検出パワーを示す。
- 理論的解析により、信号強度が適切な割合で増加する限り、次元mが増加するに従い、US手順のパワーが確率的に1に近づくことが示された。
- 本手法は未知のスパarsityパターンに対してロバストであり、非ゼロ平均差の和集合サポートの事前知識を必要としない。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。