[论文解读] Powerful genome-wide design and robust statistical inference in two-sample summary-data Mendelian randomization
本文提出了一种全基因组两样本汇总数据孟德尔随机化方法,使用超过一千个遗传工具变量和经验部分贝叶斯估计器,以增强统计功效和稳健性。通过根据工具变量的真实强度对它们进行加权,该方法提高了因果效应估计的精确性,减少了弱工具变量和多效性带来的偏倚,并显著缩小了置信区间,主要发现显示体质指数(BMI)对缺血性中风的因果优势比为1.19(95%置信区间:1.07–1.32),高密度脂蛋白胆固醇(HDL-C)对冠状动脉疾病的影响为0.78(95%置信区间:0.73–0.84)。
Two-sample summary-data Mendelian randomization (MR) has become a popular research design to estimate the causal effect of risk exposures. With the sample size of GWAS continuing to increase, it is now possible to utilize genetic instruments that are only weakly associated with the exposure. To maximize the statistical power of MR, we propose a genome-wide design where more than a thousand genetic instruments are used. For the statistical analysis, we use an empirical partially Bayes approach where instruments are weighted according to their strength, thus weak instruments bring less variation to the estimator. The estimator is highly efficient with many weak genetic instruments and is robust to balanced and/or sparse pleiotropy. We apply our method to estimate the causal effect of body mass index (BMI) and major blood lipids on cardiovascular disease outcomes and obtain substantially shorter confidence intervals. Some new and statistically significant findings are: the estimated causal odds ratio of BMI on ischemic stroke is 1.19 (95% CI: 1.07--1.32, p-value < 0.001); the estimated causal odds ratio of high-density lipoprotein cholesterol (HDL-C) on coronary artery disease (CAD) is 0.78 (95% CI 0.73--0.84, p-value < 0.001). However, the estimated effect of HDL-C becomes substantially smaller and statistically non-significant when we only use the strong instruments. By employing a genome-wide design and robust statistical methods, the statistical power of MR studies can be greatly improved. Our empirical results suggest that, even though the relationship between HDL-C and CAD appears to be highly heterogeneous, it may be too soon to completely dismiss the HDL hypothesis.
研究动机与目标
- 为克服传统两样本MR研究中仅依赖少数全基因组显著SNP而导致的统计功效有限的问题。
- 通过利用全基因组范围内的数千个遗传变异,解决MR中弱工具变量和多效性带来的挑战。
- 开发一种在平衡或多态性稀疏性及GWAS汇总数据中测量误差下仍有效的稳健、高效估计器。
- 利用真实世界GWAS数据,提高BMI和血脂对心血管疾病结局因果效应估计的精确性。
提出的方法
- 提出一种全基因组设计,使用全基因组范围内超过1,000个常见遗传变异作为工具变量。
- 引入一种经验部分贝叶斯方法,根据工具变量估计的真实强度对它们进行加权,以减少弱工具变量带来的方差。
- 将稳健调整轮廓得分方法进行适配,以校正GWAS汇总统计量中的测量误差和弱工具变量带来的偏倚。
- 采用spline-and-slab先验框架对工具变量有效性进行建模,以在不确定性下提升估计效率。
- 采用两阶段估计程序:首先估计每个工具变量的效应,然后通过考虑异方差性和多效性的收缩加权平均方法合并结果。
- 整合诊断工具,如标准化残差图和分位数-分位数图,以评估模型假设和工具变量异质性。
实验结果
研究问题
- RQ1与仅使用全基因组显著SNP相比,使用全基因组遗传工具变量集合是否能显著提升两样本MR的统计功效?
- RQ2当存在大量弱工具变量和多效性时,如何提升估计效率和稳健性?
- RQ3在考虑弱工具变量和多效性后,HDL-C对冠状动脉疾病的影响的因果效应是否仍具有显著性?
- RQ4仅限制在强工具变量上进行分析,对HDL-C对CAD的因果效应估计有何影响?
- RQ5在汇总数据MR中存在测量误差和异质性多效性时,能否实现稳健的统计推断?
主要发现
- BMI对缺血性中风的因果优势比估计为1.19(95%置信区间:1.07–1.32,p ≤ 0.001),表明存在统计显著的因果效应。
- HDL-C对冠状动脉疾病的因果优势比为0.78(95%置信区间:0.73–0.84,p ≤ 0.001),提示具有保护作用。
- 当仅限于强工具变量时,HDL-C对CAD的估计效应显著变小且不再具有显著性,凸显了纳入弱工具变量的重要性。
- 该方法获得的置信区间明显比以往MR研究更窄,显著提高了因果效应估计的精确性。
- 诊断图显示,模型假设得到合理满足,标准化残差与工具变量权重大致独立,且服从标准正态分布。
- R包mr.raps提供了复现所有结果(包括诊断)的代码,确保了研究的可重复性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。