[论文解读] Bayesian Structured Sparsity from Gaussian Fields
该论文提出了一种贝叶斯结构化稀疏模型,利用高斯过程实现基于相似性的回归系数收缩,通过椭圆切片采样实现可处理的计算。该方法通过利用特征相似性、后验包含概率和模型平均,提高了关联映射的敏感性和精确度,在基因组数据分析中优于现有方法。
Substantial research on structured sparsity has contributed to analysis of many different applications. However, there have been few Bayesian procedures among this work. Here, we develop a Bayesian model for structured sparsity that uses a Gaussian process (GP) to share parameters of the sparsity-inducing prior in proportion to feature similarity as defined by an arbitrary positive definite kernel. For linear regression, this sparsity-inducing prior on regression coefficients is a relaxation of the canonical spike-and-slab prior that flattens the mixture model into a scale mixture of normals. This prior retains the explicit posterior probability on inclusion parameters---now with GP probit prior distributions---but enables tractable computation via elliptical slice sampling for the latent Gaussian field. We motivate development of this prior using the genomic application of association mapping, or identifying genetic variants associated with a continuous trait. Our Bayesian structured sparsity model produced sparse results with substantially improved sensitivity and precision relative to comparable methods. Through simulations, we show that three properties are key to this improvement: i) modeling structure in the covariates, ii) significance testing using the posterior probabilities of inclusion, and iii) model averaging. We present results from applying this model to a large genomic dataset to demonstrate computational tractability.
研究动机与目标
- 开发一种结合高斯过程先验以通过特征相似性实现结构化稀疏性的贝叶斯框架。
- 在高维回归设置($p \gg n$)中实现可处理的计算,同时保持对特征包含的显式后验推断。
- 通过建模协变量结构并利用后验包含概率进行显著性检验,提高在基因组数据中检测真实关联的能力。
- 提供一种灵活的通用先验,可通过正定核函数适应任意预测变量之间的相似性度量。
- 在大规模基因组数据集中展示计算可处理性和相对于现有方法的统计优越性。
提出的方法
- 该方法在probit回归模型的潜变量上使用高斯过程先验,以在回归系数中诱导结构化稀疏性。
- 通过正态分布的尺度混合将spike-and-slab先验松弛为连续且可处理的形式,同时保持后验包含概率。
- 使用椭圆切片采样对潜变量高斯场进行采样,实现在高维设置下的高效后验计算。
- 通过任意正定核矩阵编码特征相似性,允许基于领域的相似性度量引导收缩。
- 应用模型平均以提高稳健性并降低特征选择中对样本偏差的敏感性。
- 该框架支持组内密集和潜在的组内稀疏稀疏性,通过修改局部正则化项实现。
实验结果
研究问题
- RQ1能否通过高斯过程利用特征相似性的贝叶斯结构化稀疏模型,提升高维回归中的敏感性和精确度?
- RQ2与标准惩罚或非贝叶斯方法相比,将潜变量高斯场用于包含概率如何影响模型选择性能?
- RQ3后验包含概率和模型平均在检测弱关联预测变量方面的统计功效提升程度如何?
- RQ4该方法能否在包含数百万个SNP的大规模基因组数据集中实现可扩展计算?
- RQ5用于度量特征相似性的核函数选择在结构化数据中恢复真实关联方面的影响力如何?
主要发现
- 与可比方法相比,所提出的模型在识别与数量性状相关的遗传变异方面显著提升了敏感性和精确度。
- 模拟结果表明,三个因素——建模协变量结构、通过后验包含概率进行显著性检验以及模型平均——是性能提升的关键。
- 该方法在大规模基因组数据集中保持了计算可处理性,使包含数万名个体和最多4000万个SNP的研究成为可能。
- 使用高斯过程先验允许灵活整合任意预测变量之间的相似性度量,增强了对领域特定结构的适应能力。
- 该模型优于强制组内密集选择的组稀疏方法,倾向于生成更具可解释性、稀疏性更好且统计功效更高的解。
- 通过基因组分窗并行化,该框架支持可扩展推理,未来还可通过多尺度或变分方法进一步扩展。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。