Skip to main content
QUICK REVIEW

[論文レビュー] Integrative High Dimensional Multiple Testing with Heterogeneity under Data Sharing Constraints

Molei Liu, Xia Yin|PubMed|Apr 2, 2020
Statistical Methods in Clinical Trials参考文献 44被引用数 11
ひとこと要約

本稿では、データ共有制約下における高次元多重仮説検定に対して、データシャーディング統合的検定手法を提案する。この手法により、研究間の不均一性を考慮しつつ、誤り発見率(FDR)の制御が可能となり、個別データの共有を必要としない。デバイアスドLASSO推定と要約統計に基づく推論を組み合わせることで、個別データメタアナリシスに匹敵する検出力が達成され、生データの共有を要しない。

ABSTRACT

Identifying informative predictors in a high dimensional regression model is a critical step for association analysis and predictive modeling. Signal detection in the high dimensional setting often fails due to the limited sample size. One approach to improving power is through meta-analyzing multiple studies which address the same scientific question. However, integrative analysis of high dimensional data from multiple studies is challenging in the presence of between-study heterogeneity. The challenge is even more pronounced with additional data sharing constraints under which only summary data can be shared across different sites. In this paper, we propose a novel data shielding integrative large-scale testing (DSILT) approach to signal detection allowing between-study heterogeneity and not requiring the sharing of individual level data. Assuming the underlying high dimensional regression models of the data differ across studies yet share similar support, the proposed method incorporates proper integrative estimation and debiasing procedures to construct test statistics for the overall effects of specific covariates. We also develop a multiple testing procedure to identify significant effects while controlling the false discovery rate (FDR) and false discovery proportion (FDP). Theoretical comparisons of the new testing procedure with the ideal individual-level meta-analysis (ILMA) approach and other distributed inference methods are investigated. Simulation studies demonstrate that the proposed testing procedure performs well in both controlling false discovery and attaining power. The new method is applied to a real example detecting interaction effects of the genetic variants for statins and obesity on the risk for type II diabetes.

研究の動機と目的

  • 標高が小さく、個別データが共有できない状況下で、高次元回帰におけるシグナル検出の課題に対処すること。
  • データ共有制約下で、誤り発見率(FDR)と誤り発見率割合(FDP)を制御する統合的多重仮説検定手法を開発すること。
  • 個別データを必要とせず、研究間の不均一性を高次元モデルで適切に扱い、プライバシー準拠を確保すること。
  • 現実的なデータ共有制限下で、要約統計のみを用いて複数の研究における共変量効果の同時推論を可能にすること。
  • 理想の個別データメタアナリシスと分散型推論のギャップを埋め、通信コストを最小限に抑えつつ、同等の検出力を達成すること。

提案手法

  • 2段階の統合的推定手順を提案:まず、各拠点でローカルなデバイアスドLASSO推定器を、ローカルデータと要約統計のみを用いて計算する。
  • 次に、分析センターが、グループ構造に基づく切断を用いて、これらのローカルでデバイアスされた推定器を統合し、グローバル統合推定器を形成する。
  • 研究間の不均一性を考慮するデバイアスフレームワークを用いて、各共変量の検定統計量を構築し、弱スパarsity仮定の下で漸近正規性を保証する。
  • Benjamini-Hochberg手順に基づく多重仮説検定手順を用いて、すべての $ p $ 個の共変量におけるFDRとFDPを制御する。
  • デバイアスステップにおける適切な分散推定を可能にするために、各データ拠点から分析センターへヘッセ行列を転送する。
  • データ共有制約下でも、$ p $ と同時に $ M $(研究数)が発散する場合でも理論的保証を維持する。

実験結果

リサーチクエスチョン

  • RQ1個別データが研究間で共有されない状況下でも、高次元多重仮説検定において高い統計的検出力を達成できるか。
  • RQ2研究間の不均一性とデータ共有制約下で、統合的解析における誤り発見率と誤り発見率割合をどのように制御できるか。
  • RQ3提案手法と理想の個別データメタアナリシスとの間の理論的関係(検出力と誤差制御の観点から)は何か。
  • RQ4要約統計のみを用いる場合でも、個別データ手法と同等のスパarsity仮定を維持できるか。
  • RQ5本手法の通信複雑度は何か。また、統計的効率を損なわずにこれを軽減できるか。

主な発見

  • 提案手法は、弱スパarsity仮定の下で、漸近的に誤り発見率と誤り発見率割合を制御する。
  • 本手法は、理想の個別データメタアナリシスに匹敵する統計的検出力を達成しており、ワンショット法よりも検出力と頑健性に優れる。
  • 本手法のスパarsity仮定は理想手法と同等であるが、ワンショット法が要請するものよりも厳密に弱い。
  • ワンショット法と比較して、ヘッセ行列の送信のための追加通信ラウンド1回のみを要するため、実世界の応用において実用的である。
  • スタチン-遺伝子相互作用と2型糖尿病に関する実データ応用では、本手法により90%信頼区間をデバイアス手順で推定した上で、5つの有意な相互作用効果を検出できた。
  • 理論的解析により、帰無仮説下で検定統計量が漸近的に正規分布に従うことが確認され、有効な同時推論が可能であることが示された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。