[论文解读] Inference After Selecting Plausibly Valid Instruments with Application to Mendelian Randomization
本文提出了一种条件推断框架,通过选择性推断校正孟德尔随机化(MR)中数据驱动的工具变量选择偏差。通过针对所选工具变量(特别是sisVIVE程序)进行条件处理,该方法推导出有效的零分布和置信区间,实现了名义上的覆盖水平,而标准方法忽略选择过程,导致反保守推断。
Mendelian randomization (MR) is a popular method in genetic epidemiology to estimate the effect of an exposure on an outcome by using genetic instruments. These instruments are often selected from a combination of prior knowledge from genome wide association studies (GWAS) and data-driven instrument selection procedures or tests. Unfortunately, when testing for the exposure effect, the instrument selection process done a priori is not accounted for. This paper studies and highlights the bias resulting from not accounting for the instrument selection process by focusing on a recent data-driven instrument selection procedure, sisVIVE, as an example. We introduce a conditional inference approach that conditions on the instrument selection done a priori and leverage recent advances in selective inference to derive conditional null distributions of popular test statistics for the exposure effect in MR. The null distributions can be characterized with individual-level or summary-level data in MR. We show that our conditional confidence intervals derived from conditional null distributions attain the desired nominal level while typical confidence intervals computed in MR do not. We conclude by reanalyzing the effect of BMI on diastolic blood pressure using summary-level data from the UKBiobank that accounts for instrument selection.
研究动机与目标
- 解决在孟德尔随机化(MR)分析中忽略工具变量选择过程所引入的偏差。
- 开发一种条件推断方法,以在MR环境中考虑基于数据的工具变量选择(如sisVIVE)。
- 推导出对所选工具变量进行条件处理的暴露效应的有效零分布和置信区间。
- 证明标准MR推断由于选择偏差无法维持名义覆盖水平,而所提方法可实现该目标。
- 将该方法应用于英国生物样本库的真实世界汇总水平数据,重新分析体质指数(BMI)对舒张压的因果效应。
提出的方法
- 使用选择性推断,对通过数据驱动程序(如sisVIVE)选择特定工具变量集合的事件进行条件处理。
- 在选定的工具变量集合下,推导检验统计量(如Anderson-Rubin、两阶段最小二乘法)的条件零分布。
- 将该框架应用于个体水平和汇总水平的MR数据,使其在大型遗传学研究中具有实际应用价值。
- 利用近期在选择后推断方面的进展,调整因工具变量基于其强度和有效性检验而被选择的事实。
- 构建基于选择事件的置信区间,确保正确的频率覆盖。
- 使用随机化程序处理工具变量有效性中的模糊性,相比联合方法提高了稳健性和统计功效。
实验结果
研究问题
- RQ1在MR中,为何忽略基于数据的工具变量选择会导致无效推断?
- RQ2选择性推断能否被调整以校正基于工具变量强度和多效性检验选择工具时的偏差?
- RQ3基于条件零分布推导的置信区间在有限样本中是否能维持名义覆盖水平?
- RQ4与现有方法(如Kang等,2015年)相比,所提方法在区间宽度和覆盖水平方面表现如何?
- RQ5工具变量选择对MR研究中的第一类错误率和统计功效有何影响?
主要发现
- 忽略工具变量选择的标准MR置信区间无法实现名义上的95%覆盖水平,导致反保守推断。
- 所提出的条件推断框架通过基于所选工具变量进行条件处理,生成实现目标名义覆盖水平的置信区间。
- 在英国生物样本库的再分析中,BMI对舒张压的条件置信区间比标准区间更窄且更准确。
- 当 $U=2$ 时,使用TSLS的有选择性置信区间为 $(0.034, 0.058)$,而Kang等(2015年)的联合方法得到 $(0.031, 0.059)$,表明结果相近但略为精确。
- 与联合方法相比,该方法通过聚焦于实际选择结果而非所有可能的无效工具组合,提高了统计功效并缩小了置信区间宽度。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。