Skip to main content
QUICK REVIEW

[论文解读] Bayesian Sparse Mediation Analysis with Targeted Penalization of Natural Indirect Effects

Yanyi Song, Xiang Zhou|arXiv (Cornell University)|Aug 14, 2020
Advanced Causal Inference Techniques参考文献 12被引用 4
一句话总结

本文提出两种新颖的贝叶斯方法——GMM 和 PTG,用于高维中介分析,通过直接惩罚自然中介效应(NIE)来识别活跃中介。通过使用结构化先验联合建模暴露-中介和中介-结局效应,该方法提升了选择与估计的准确性,在模拟研究中相较于现有方法最高实现30%的效能提升,并在 MESA 和 LIFECODES 队列中识别出具有生物学意义的中介因子。

ABSTRACT

Causal mediation analysis aims to characterize an exposure's effect on an outcome and quantify the indirect effect that acts through a given mediator or a group of mediators of interest. With the increasing availability of measurements on a large number of potential mediators, like the epigenome or the microbiome, new statistical methods are needed to simultaneously accommodate high-dimensional mediators while directly target penalization of the natural indirect effect (NIE) for active mediator identification. Here, we develop two novel prior models for identification of active mediators in high-dimensional mediation analysis through penalizing NIEs in a Bayesian paradigm. Both methods specify a joint prior distribution on the exposure-mediator effect and mediator-outcome effect with either (a) a four-component Gaussian mixture prior or (b) a product threshold Gaussian prior. By jointly modeling the two parameters that contribute to the NIE, the proposed methods enable penalization on their product in a targeted way. Resultant inference can take into account the four-component composite structure underlying the NIE. We show through simulations that the proposed methods improve both selection and estimation accuracy compared to other competing methods. We applied our methods for an in-depth analysis of two ongoing epidemiologic studies: the Multi-Ethnic Study of Atherosclerosis (MESA) and the LIFECODES birth cohort. The identified active mediators in both studies reveal important biological pathways for understanding disease mechanisms.

研究动机与目标

  • 解决在测量数千种潜在中介因子(如 DNAm、生物标志物)的高维环境中识别活跃中介因子的挑战。
  • 克服单变量中介分析和标准正则化方法的局限性,后者分别惩罚路径系数而未针对中介效应本身。
  • 开发一种贝叶斯框架,联合建模暴露-中介和中介-结局效应,实现对中介效应乘积(NIE)的直接惩罚,以提升中介因子选择能力。
  • 在复杂高维组学数据中提升识别具有强间接效应中介因子的统计效能,并减少估计偏差。

提出的方法

  • 提出两种联合先验模型:四分量高斯混合先验(GMM)和乘积阈值高斯先验(PTG),用于联合建模暴露-中介和中介-结局效应。
  • 通过反映中介因子四类复合结构的结构化先验,对自然中介效应(NIE)实施针对性惩罚,NIE 定义为暴露-中介与中介-结局路径系数的乘积。
  • 采用中位数包含概率(PIPs)为 0.5 作为识别活跃中介因子的选择标准,并通过模拟研究验证其对错误发现率的控制能力。
  • 通过层次化先验实现收缩与变量选择,使高维环境中能够同时完成活跃中介因子的估计与选择。
  • 将方法应用于 MESA 和 LIFECODES 队列的真实数据,分析 DNAm 和生物标志物中介因子与健康结局的关系。
  • 通过模拟研究将性能与竞争方法(包括 Lasso、SCAD 和贝叶斯变量选择)进行比较,重点关注 NIE 的选择与估计准确性。

实验结果

研究问题

  • RQ1基于针对自然中介效应的贝叶斯联合建模与针对性惩罚,能否提升高维中介分析中活跃中介因子的识别能力?
  • RQ2与现有惩罚方法及贝叶斯方法相比,所提出方法在识别真正非零中介因子方面的效能与准确性如何?
  • RQ3在将该方法应用于 MESA 和 LIFECODES 的真实高维组学数据时,揭示了哪些生物通路?
  • RQ4中位数包含概率(PIP = 0.5)作为选择阈值,在此情境下对错误发现率的控制效果如何?
  • RQ5中介因子之间的相关性在多大程度上影响所提方法的性能?未来模型如何更好地整合这些相关性?

主要发现

  • 所提出的 GMM 和 PTG 方法在模拟研究中识别真正非零中介因子的统计效能最高比竞争方法高出30%。
  • 在 MESA 队列中,该方法识别出8–10个关键 DNAm 位点及其附近基因(如 NFE2L1、PTK2 和 CREB1),介导了社区社会经济地位对 BMI 的影响。
  • 在 LIFECODES 队列中,该方法检测到 12(13)-EpoME 和 9-oxoODE 为显著中介因子,连接孕期邻苯二甲酸盐暴露与妊娠龄,提示其在氧化应激与炎症通路中的作用。
  • 在选择准确性与估计偏差减少方面,该方法优于现有贝叶斯与频率学派方法,尤其在间接效应中等至较强时表现更优。
  • 将中位数包含概率(PIP = 0.5)作为选择阈值,有效控制了错误发现率,该结论通过模拟实验得到验证。
  • 该方法揭示了具有生物学合理性的中介通路,如 DNAm 介导的代谢与炎症相关表型效应,支持其在生物社会学与流行病学研究中的应用价值。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。