[论文解读] Integrative High Dimensional Multiple Testing with Heterogeneity under Data Sharing Constraints
本文提出了一种在数据共享限制下针对高维多重检验的数据屏蔽整合测试方法,可在考虑研究间异质性的同时控制错误发现率。通过结合去偏LASSO估计与基于汇总统计量的推断,该方法在无需共享原始数据的情况下,实现了与个体水平元分析相当的检验效能。
Identifying informative predictors in a high dimensional regression model is a critical step for association analysis and predictive modeling. Signal detection in the high dimensional setting often fails due to the limited sample size. One approach to improving power is through meta-analyzing multiple studies which address the same scientific question. However, integrative analysis of high dimensional data from multiple studies is challenging in the presence of between-study heterogeneity. The challenge is even more pronounced with additional data sharing constraints under which only summary data can be shared across different sites. In this paper, we propose a novel data shielding integrative large-scale testing (DSILT) approach to signal detection allowing between-study heterogeneity and not requiring the sharing of individual level data. Assuming the underlying high dimensional regression models of the data differ across studies yet share similar support, the proposed method incorporates proper integrative estimation and debiasing procedures to construct test statistics for the overall effects of specific covariates. We also develop a multiple testing procedure to identify significant effects while controlling the false discovery rate (FDR) and false discovery proportion (FDP). Theoretical comparisons of the new testing procedure with the ideal individual-level meta-analysis (ILMA) approach and other distributed inference methods are investigated. Simulation studies demonstrate that the proposed testing procedure performs well in both controlling false discovery and attaining power. The new method is applied to a real example detecting interaction effects of the genetic variants for statins and obesity on the risk for type II diabetes.
研究动机与目标
- 解决在样本量有限且无法共享个体水平数据时,高维回归中信号检测的挑战。
- 开发一种在数据共享限制下控制错误发现率(FDR)和错误发现比例(FDP)的整合多重检验方法。
- 在不依赖个体水平数据的前提下,考虑高维模型中的研究间异质性,确保符合隐私合规要求。
- 仅使用汇总统计量,在现实数据共享限制下,实现对多个研究中协变量效应的联合推断。
- 通过实现与理想个体水平元分析相当的检验效能,同时保持最低通信开销,弥合理想个体水平元分析与分布式推断之间的差距。
提出的方法
- 提出两步整合估计程序:首先,在每个站点仅使用本地数据和汇总统计量计算局部去偏LASSO估计量。
- 其次,分析中心使用分组结构截断方法聚合这些局部去偏估计量,形成全局整合估计量。
- 基于去偏框架构建每个协变量的检验统计量,以考虑研究间的异质性,确保在弱稀疏性假设下的渐近正态性。
- 采用基于Benjamini-Hochberg程序的多重检验方法,控制所有 $ p $ 个协变量的FDR和FDP。
- 从每个数据站点向分析中心传输Hessian矩阵,以在去偏步骤中实现正确的方差估计。
- 允许研究数量 $ M $ 随 $ p $ 发散,同时在数据共享限制下保持理论保证。
实验结果
研究问题
- RQ1当各研究间不共享个体水平数据时,我们能否在高维多重检验中实现高统计效能?
- RQ2在研究间存在异质性且受数据共享限制的整合分析中,如何控制错误发现率和错误发现比例?
- RQ3所提出方法与理想个体水平元分析在效能和误差控制方面的理论关系是什么?
- RQ4在仅使用汇总统计量的前提下,我们能否维持与个体水平方法相当的稀疏性假设?
- RQ5该方法的通信复杂度是多少?是否可在不牺牲统计效率的前提下降低?
主要发现
- 在弱稀疏性假设下,所提方法可渐近控制错误发现率和错误发现比例。
- 该方法的统计效能与理想个体水平元分析相当,且在效能和稳健性方面均优于单次通信方法。
- 所提方法的稀疏性假设与理想方法等价,但比单次通信方法所需假设更弱。
- 与单次通信方法相比,该方法仅需额外一轮通信(用于传输Hessian矩阵),使其在实际应用中更具可行性。
- 在一项关于他汀类药物-基因互作与2型糖尿病的真实数据应用中,该方法成功检测出5个显著的互作效应,并通过所提出的去偏程序估计了90%的置信区间。
- 理论分析证实,在原假设下,检验统计量渐近服从正态分布,从而支持有效的联合推断。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。