[论文解读] Statistical Methods for cis-Mendelian Randomization
本文通过使用单个基因区域的遗传变异评估了顺式孟德尔随机化中变量选择与因果效应估计方法,特别关注蛋白质表达作为风险因素的情形。研究发现,当工具变量较弱时,因子分析和贝叶斯变量选择优于简单修剪法,从而确保更可靠的因果推断。
Mendelian randomization is the use of genetic variants to assess the existence of a causal relationship between a risk factor and an outcome of interest. In this paper we focus on Mendelian randomization analyses with many correlated variants from a single gene region, and particularly on cis-Mendelian randomization studies which uses protein expression as a risk factor. Such studies must rely on a small, curated set of variants from the studied region; using all variants in the region requires inverting an ill-conditioned genetic correlation matrix and results in numerically unstable causal effect estimates. We review methods for variable selection and causal effect estimation in cis-Mendelian randomization, ranging from stepwise pruning and conditional analysis to principal components analysis, factor analysis and Bayesian variable selection. In a simulation study, we show that the various methods have a comparable performance in analyses with large sample sizes and strong genetic instruments. However, when weak instrument bias is suspected, factor analysis and Bayesian variable selection produce more reliable inference than simple pruning approaches, which are often used in practice. We conclude by examining two case studies, assessing the effects of LDL-cholesterol and serum testosterone on coronary heart disease risk using variants in the HMGCR and SHBG gene regions respectively.
研究动机与目标
- 为解决因使用基因区域中所有相关变异而导致顺式孟德尔随机化中数值不稳定的問題。
- 评估在不同遗传工具强度下,多种变量选择技术(如逐步修剪、条件分析、主成分分析、因子分析和贝叶斯变量选择)的性能。
- 在弱工具变量存在的情况下,提高因果效应估计的可靠性,因为标准修剪方法可能产生偏倚。
- 为研究人员在以蛋白质表达作为风险因素的顺式孟德尔随机化研究中提供方法指导。
- 通过模拟研究和对低密度脂蛋白胆固醇及睾酮对冠心病影响的真实案例研究,验证方法性能。
提出的方法
- 使用单个基因区域的遗传变异精选子集,以避免遗传相关矩阵病态倒置的问题。
- 应用逐步修剪法,基于相关性阈值迭代剔除变异。
- 采用条件分析法,在选择变异时对先前选定的变异进行调整。
- 利用主成分分析(PCA)将相关变异总结为正交成分。
- 应用因子分析提取代表潜在遗传结构的潜在因子。
- 使用贝叶斯变量选择,通过后验包含概率概率性地识别最相关的变异。
实验结果
研究问题
- RQ1当遗传工具变量较强或较弱时,不同变量选择方法在顺式孟德尔随机化中的表现如何?
- RQ2当遗传相关矩阵病态时,哪种方法能产生最可靠的因果效应估计?
- RQ3在弱工具变量条件下,因子分析和贝叶斯变量选择与标准修剪法相比,在偏倚和精度方面表现如何?
- RQ4方法选择对真实世界顺式孟德尔随机化研究中因果推断的影响是什么?
- RQ5这些方法能否有效应用于评估蛋白质表达对疾病结局的因果效应?
主要发现
- 在大样本量且工具变量较强时,所有方法(包括修剪、PCA、因子分析和贝叶斯选择)在因果效应估计方面表现相当。
- 当存在弱工具变量偏倚时,因子分析和贝叶斯变量选择比简单修剪法提供更可靠的推断。
- 实践中常用的修剪方法在工具变量较弱时易产生偏倚,导致第一类错误率膨胀。
- 因子分析和贝叶斯变量选择能更好地捕捉变异间多变量相关结构,从而提升估计稳定性。
- 模拟研究证实,这些高级方法在弱工具变量条件下可降低偏倚并提高置信区间覆盖概率。
- 对HMGCR和SHBG基因区域的真实案例研究展示了这些方法在评估低密度脂蛋白胆固醇和睾酮对冠心病因果效应方面的实际应用价值。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。