Skip to main content
QUICK REVIEW

[论文解读] Modular Bayes screening for high-dimensional predictors

Yuhan Chen, David B. Dunson|arXiv (Cornell University)|Mar 29, 2017
Bayesian Methods and Mixture Models参考文献 21被引用 4
一句话总结

本文提出 MOBS(模块化贝叶斯筛选),一种用于高维预测变量的贝叶斯非参数筛选方法,通过模块化设计将响应变量边际分布建模与预测变量特异性效应建模解耦。通过在响应变量上应用狄利克雷过程混合模型,并条件计算每个预测变量的后验包含概率,MOBS 在跨检验中实现信息借用,同时避免了不确定性低估问题,在包含 3800 万个 SNP 的顺式-eQTL 数据集中,其在检测超出均值差异的复杂分布变化方面优于现有方法。

ABSTRACT

With the routine collection of massive-dimensional predictors in many application areas, screening methods that rapidly identify a small subset of promising predictors have become commonplace. We propose a new MOdular Bayes Screening (MOBS) approach, which involves several novel characteristics that can potentially lead to improved performance. MOBS first applies a Bayesian mixture model to the marginal distribution of the response, obtaining posterior samples of mixture weights, cluster-specific parameters, and cluster allocations for each subject. Hypothesis tests are then introduced, corresponding to whether or not to include a given predictor, with posterior probabilities for each hypothesis available analytically conditionally on unknowns sampled in the first stage and tuning parameters controlling borrowing of information across tests. By marginalizing over the first stage posterior samples, we avoid under-estimation of uncertainty typical of two-stage methods. We greatly simplify the model specification and reduce computational complexity by using {\em modularization}. We provide basic theoretical support for this approach, and illustrate excellent performance relative to competitors in simulation studies and the ability to capture complex shifts beyond simple differences in means. The method is illustrated with applications to genomics by using a very high-dimensional cis-eQTL dataset with roughly 38 million SNPs.

研究动机与目标

  • 解决现有高维筛选方法依赖强参数假设或无法在检验间共享信息的局限性。
  • 开发一种可扩展、计算高效的筛选方法,通过避免两阶段估计的缺陷来保持不确定性量化。
  • 在 p 大 n 小的场景下,利用非参数贝叶斯模型检测超出简单均值差异的复杂分布变化。
  • 通过模块化建模,提供一种适用于多种数据类型(包括分类、连续及复杂对象数据)的灵活框架。

提出的方法

  • MOBS 采用模块化贝叶斯方法,首先对响应变量 y 的边际分布拟合狄利克雷过程混合模型,以获得混合权重、聚类参数和聚类分配的后验样本。
  • 在这些后验样本的条件下,利用分层先验结构计算每个预测变量的变量包含后验概率,从而实现在检验间的知识共享。
  • 通过在第一阶段后验样本上进行积分,对不确定性进行边际化处理,避免了两阶段方法中常见的方差低估问题。
  • 通过模块化设计将响应模型与预测变量模型解耦,提升了在高维设置下的计算可行性与鲁棒性。
  • 对于连续或混合预测变量,采用基于节点的离散化方法与核插值技术,以估计条件密度 f(y|x_j)。
  • 只要基准响应模型存在合适的 MCMC 算法,该方法可扩展至多变量及复杂结果类型。

实验结果

研究问题

  • RQ1模块化贝叶斯非参数方法是否能通过在检验间共享信息而无需强参数假设,从而提升高维设置下的筛选性能?
  • RQ2MOBS 在检测超出均值差异的复杂分布变化方面,与现有筛选方法相比表现如何?
  • RQ3MOBS 在多大程度上保持了不确定性量化,并避免了两阶段方法中常见的方差低估问题?
  • RQ4MOBS 是否能有效识别超大维度基因组数据(如约 3800 万个 SNP 的顺式-eQTL 数据集)中的生物相关 SNP?

主要发现

  • MOBS 识别出的 SNP 总绝对均值距离显著更低(1.12),优于 DCS(1.88)、SIS(1.91)、SIRS(1.86)和 JYL(1.64),表明其能捕捉超出均值差异的复杂分布变化。
  • 在 50 个 SNP 子集上的预测性能中,MOBS-L 的平均 MSE 为 0.850,与 SIS-L(0.880)、DCS-L(0.882)、SIRS-L(0.886)、FUSEDK-L(0.847)和 JYL-L(0.852)相比表现相当或更优。
  • MOBS 选择的 SNP 集合与 SIS、DCS 或 SIRS 主要识别的 SNP 不同,后者多集中于检测均值变化,表明 MOBS 具备检测新型复杂效应的能力。
  • 该方法成功处理了约 3800 万个 SNP 的顺式-eQTL 数据集,展示了在超大维度设置下的可扩展性。
  • MOBS 在模拟与真实数据中均表现优异,能有效检测非基于均值的变化,同时保持计算效率与不确定性校准。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。