[论文解读] Semi-parametric Bayesian variable selection for gene-environment interactions
本文提出了一种新颖的半参数贝叶斯变量选择模型,能够同时识别高维基因组数据中的线性和非线性基因-环境(G×E)互作效应。通过使用B样条函数建模非线性效应,并在个体和组别两个层次上采用层次化稀疏-滑动先验,该方法实现了自动结构识别——区分仅存在主效应、存在G×E互作效应以及无遗传效应三种情形,在模拟和真实数据分析中均表现出优越性能。
Many complex diseases are known to be affected by the interactions between genetic variants and environmental exposures beyond the main genetic and environmental effects. Study of gene-environment (G$ imes$E) interactions is important for elucidating the disease etiology. Existing Bayesian methods for G$ imes$E interaction studies are challenged by the high-dimensional nature of the study and the complexity of environmental influences. Many studies have shown the advantages of penalization methods in detecting G$ imes$E interactions in "large p, small n" settings. However, Bayesian variable selection, which can provide fresh insight into G$ imes$E study, has not been widely examined. We propose a novel and powerful semi-parametric Bayesian variable selection model that can investigate linear and nonlinear G$ imes$E interactions simultaneously. Furthermore, the proposed method can conduct structural identification by distinguishing nonlinear interactions from main-effects-only case within the Bayesian framework. Spike and slab priors are incorporated on both individual and group levels to identify the sparse main and interaction effects. The proposed method conducts Bayesian variable selection more efficiently than existing methods. Simulation shows that the proposed model outperforms competing alternatives in terms of both identification and prediction. The proposed Bayesian method leads to the identification of main and interaction effects with important implications in a high-throughput profiling study with high-dimensional SNP data.
研究动机与目标
- 为解决现有贝叶斯方法在高维'大p,小n'设定下检测复杂G×E互作效应的局限性。
- 开发一种能够同时识别线性和非线性G×E互作效应的方法。
- 在贝叶斯框架内实现结构识别,区分仅存在主效应、存在G×E互作效应以及无遗传效应三种情形。
- 与现有惩罚法和贝叶斯方法相比,提升变量选择效率和预测准确性。
提出的方法
- 使用B样条基展开来建模非线性G×E互作效应,以灵活捕捉基因-环境效应中的非线性模式。
- 在个体和组别两个层次上实施具有稀疏-滑动先验的层次化贝叶斯模型,以实现对重要效应的稀疏选择。
- 对一组样条系数使用多变量拉普拉斯先验,对单个系数使用单变量拉普拉斯先验,以诱导收缩与选择。
- 采用高效的吉布斯采样器,并通过C++加速核心模块,实现在高维设定下的快速MCMC计算。
- 通过样条基变换实现对系数函数中变化、非零常数和零系数的自动结构识别。
- 在建模框架中整合临床协变量,并控制混杂因素,以提高估计准确性。
实验结果
研究问题
- RQ1贝叶斯变量选择方法是否能有效检测高维基因组数据中的线性和非线性基因-环境互作效应?
- RQ2所提出的方法在统一框架下能否良好地区分仅存在主效应、存在G×E互作效应以及无遗传效应的情形?
- RQ3与参数化替代方法相比,通过B样条实现的半参数建模是否能提升检测效能和预测准确性?
- RQ4层次化稀疏-滑动先验结构在高维G×E互作研究中在多大程度上提升了变量选择效率?
- RQ5与现有惩罚法和贝叶斯方法相比,该方法在真实世界高通量SNP数据中的表现如何?
主要发现
- 在多种模拟情景下,所提出方法在变量识别和预测准确性方面均优于竞争方法。
- 在Nurses' Health Study数据中,该方法成功检测到非线性G×E互作效应,图1显示线性互作假设存在明显违反。
- 在真实数据应用中,该方法识别出与rs1106380、rs10999234和rs796945等SNP相关的显著G×E互作效应,可信区间支持其生物学相关性。
- 层次化稀疏-滑动先验结构在个体和组别两个层次上均实现了对重要效应的有效选择,降低了假阳性率。
- MCMC算法实现了快速收敛和计算效率,核心模块采用C++实现,具备良好的性能可扩展性。
- 该方法在区分仅存在主效应和真实G×E互作效应方面表现出稳健性,证实了其结构识别能力。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。