Skip to main content
QUICK REVIEW

[论文解读] Scalable Approximations of Marginal Posteriors in Variable Selection

Willem van den Boom, Galen Reeves|arXiv (Cornell University)|Jun 22, 2015
Bayesian Methods and Mixture Models参考文献 25被引用 7
一句话总结

本文提出一种可扩展的框架,结合近似消息传递(AMP)与贝叶斯压缩回归(BCR),以高效近似高维贝叶斯变量选择中边际后验包含概率,采用稀疏-斯拉布先验。该方法通过近似旋转数据的后验预测分布,实现对变量显著性的精确、快速估计——这对科学推断至关重要——从而在神经影像等大规模场景中实现可靠的不确定性量化。

ABSTRACT

In many contexts, there is interest in selecting the most important variables from a very large collection, commonly referred to as support recovery or variable, feature or subset selection. There is an enormous literature proposing a rich variety of algorithms. In scientific applications, it is of crucial importance to quantify uncertainty in variable selection, providing measures of statistical significance for each variable. The overwhelming majority of algorithms fail to produce such measures. This has led to a focus in the scientific literature on independent screening methods, which examine each variable in isolation, obtaining p-values measuring the significance of marginal associations. Bayesian methods provide an alternative, with marginal inclusion probabilities used in place of p-values. Bayesian variable selection has advantages, but is impractical computationally beyond small problems. In this article, we show that approximate message passing (AMP) and Bayesian compressed regression (BCR) can be used to rapidly obtain accurate approximations to marginal inclusion probabilities in high-dimensional variable selection. Theoretical support is provided, simulation studies are conducted to assess performance, and the method is applied to a study relating brain networks to creative reasoning.

研究动机与目标

  • 解决高维变量选择中缺乏不确定性量化的问题,其中大多数快速方法无法提供统计显著性度量。
  • 克服在大规模问题(如p > 1000)中使用MCMC进行精确贝叶斯变量选择的计算不可行性。
  • 开发一种通用且可扩展的框架,用于近似非高斯先验下贝叶斯线性回归中的边际后部分布。
  • 为每个预测变量提供精确的后验包含概率(PIPs),作为科学推断中p值的贝叶斯对应物。
  • 在模拟数据和具有复杂、重尾及病态特征的真实神经影像数据上,展示该方法的稳健性与准确性。

提出的方法

  • 通过基于旋转的变换,将感兴趣的参数与干扰系数解耦,将高维后部分布简化为标量问题。
  • 使用最先进的方法(AMP与BCR)近似旋转数据的后验预测分布。
  • 应用稀疏-斯拉布先验来建模变量包含性,其中每个系数要么为零(稀疏),要么服从正态分布(斯拉布)。
  • 通过包含与排除模型证据的比值,利用近似后的后验预测密度来估计边际后验包含概率。
  • 通过经验贝叶斯或层次先验方法,对未知误差方差σ²和包含概率λ进行积分处理。
  • 利用在大p条件下后验预测近似具有渐近正态性的特性,即使完整后验分布非高斯。

实验结果

研究问题

  • RQ1AMP与BCR能否在高维贝叶斯变量选择中,对边际后验包含概率提供准确且可扩展的近似?
  • RQ2在模拟设置中,这些近似在不同相关结构(ρ)和信噪比(SNR)下的表现如何?
  • RQ3当σ²与λ未知且必须从数据中估计时,该框架能否可靠地估计PIPs?
  • RQ4与精确MCMC相比,这些近似在准确性和计算效率方面表现如何?
  • RQ5该方法能否在具有高维、离散且重尾特征的真实神经影像数据中,识别出具有生物学意义的脑网络连接?

主要发现

  • 在多种模拟设置下,包括高相关性(ρ = 0.8)和低信噪比(SNR = 1)时,AMP与BCR均能对后验包含概率提供高度准确的近似。
  • 该方法在保持与MCMC方法相当的PIPs估计准确性的同时,实现了数量级的加速。
  • 即使σ²与λ未知,AMP与BCR仍能成功估计PIPs,200次模拟的箱线图显示估计结果一致且具有意义。
  • 在真实脑连接性研究中(n=113,p=1802),该方法识别出跨半球连接为高度可能,与先前神经科学发现一致。
  • BCR与AMP在边优先级排序上表现相似但略有不同,两种方法均恢复了已知的连接模式——例如,从节点18L到3R、18R和20R的连接,尽管数据具有重尾和病态特性。
  • AMP在高相关性条件下略显激进,更倾向于将PIPs推向0或1,但两种方法均保持稳定且信息丰富。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。