Skip to main content
QUICK REVIEW

[論文レビュー] Fast and Powerful Conditional Randomization Testing via Distillation

Molei Liu, Eugene Katsevich|arXiv (Cornell University)|Jun 6, 2020
Statistical Methods in Clinical Trials参考文献 48被引用数 15
ひとこと要約

本稿では、計算効率の高い手法として、機械学習を活用した強力な条件付き独立性検定を実現しつつ、正確な第一種誤り率の制御を維持する、蒸留された条件付きランダマイゼーション検定(CRT)を提案する。複雑なモデルを軽量な代替モデルに縮約し、スクリーニングや計算再利用の技術を用いることで、フルCRTとほぼ同等の検出力を持つが、計算時間が桁違いに速くなる。この手法により、乳がん遺伝子解析を含む大規模データセットへの実用的応用が可能になる。

ABSTRACT

We consider the problem of conditional independence testing: given a response Y and covariates (X,Z), we test the null hypothesis that Y is independent of X given Z. The conditional randomization test (CRT) was recently proposed as a way to use distributional information about X|Z to exactly (non-asymptotically) control Type-I error using any test statistic in any dimensionality without assuming anything about Y|(X,Z). This flexibility in principle allows one to derive powerful test statistics from complex prediction algorithms while maintaining statistical validity. Yet the direct use of such advanced test statistics in the CRT is prohibitively computationally expensive, especially with multiple testing, due to the CRT's requirement to recompute the test statistic many times on resampled data. We propose the distilled CRT, a novel approach to using state-of-the-art machine learning algorithms in the CRT while drastically reducing the number of times those algorithms need to be run, thereby taking advantage of their power and the CRT's statistical guarantees without suffering the usual computational expense. In addition to distillation, we propose a number of other tricks like screening and recycling computations to further speed up the CRT without sacrificing its high power and exact validity. Indeed, we show in simulations that all our proposals combined lead to a test that has similar power to the most powerful existing CRT implementations but requires orders of magnitude less computation, making it a practical tool even for large data sets. We demonstrate these benefits on a breast cancer dataset by identifying biomarkers related to cancer stage.

研究の動機と目的

  • 条件付きランダマイゼーション検定(CRT)において、複数の仮説検定を伴う場合に繰り返し複雑な機械学習モデルを実行する高コストを軽減すること。
  • フルモデルの再訓練回数を大幅に削減しながらも、CRTが保証する正確な第一種誤り率と高い検出力を維持すること。
  • 標準的なCRTが非現実的となる高次元設定において、強力な機械学習ベースの検定統計量の実用的応用を可能にすること。
  • スケーラブルな条件付き独立性検定を実現するため、モデル蒸留、スクリーニング、計算再利用を統合したフレームワークの構築

提案手法

  • 元のデータ上で訓練された複雑なベースモデル(教師)の予測を模倣するように、軽量な代替モデル(生徒)を訓練する、蒸留されたCRTを提案する。
  • リサンプリングされたデータ上で、蒸留モデルを用いて検定統計量を計算することで、各リサンプリングに対してフルモデルを再訓練する必要を減らす。
  • 不要な共変量を事前にフィルタリングするスクリーニングを適用し、検定数と計算負荷を削減する。
  • 類似したデータに対して同じモデルを繰り返し評価するのを避けるために、リサンプリング間で中間計算を再利用する。
  • 蒸留モデル出力にロジスティック回帰を適用する、リサンプリングを不要とするCRTの変種を採用し、推論をさらに高速化する。
  • すべての技術を統合したパイプラインを構築し、正確な妥当性と高い検出力を維持しながら、実行時間を最小限に抑える。

実験結果

リサーチクエスチョン

  • RQ1リサンプリングデータ上で複雑なモデルを実行する負担を軽減しながら、CRTの正確な第一種誤り率の制御を維持できるか?
  • RQ2モデル蒸留は、条件付き独立性検定における機械学習ベースの検定統計量の検出力をどの程度保持できるか?
  • RQ3スクリーニングと計算再利用を組み合わせることで、妥当性や検出力に損なわれることなく、CRTの高速化はどの程度達成できるか?
  • RQ4実世界のデータにおいて、既存手法(oCRT、HRT、knockoffs)と比較して、速度、検出力、および誤発見率制御の観点から、蒸留CRTはどのように優れているか?

主な発見

  • 蒸留されたCRTは、複雑なモデルを用いても、最も強力な既存のCRT実装と同等の検出力を達成した。
  • 標準的なCRTと比較して、計算時間を桁違いに短縮し、大規模データセットへの適用が現実可能になった。
  • 乳がんデータセットでは、FDR制御下でFBXW7、GPS2、RUNX1といった既知のバイオマーカーが、高い検出頻度で同定された。
  • スクリーニングを施したdI CRT(蒸留+スクリーニング)は、α=0.1の水準で、上位遺伝子の100%をFDR制御下で検出でき、oCRT や HRT よりもキーリードジーンの検出頻度が優れていた。
  • すべてのシミュレーションおよび実データ実験において、正確な第一種誤り率の制御が維持された。統計的厳密性が裏付けられた。
  • 蒸留、スクリーニング、計算再利用の組み合わせにより、実用的に100倍~1000倍の高速化が達成され、検出力の損失なしに実現した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。