Skip to main content
QUICK REVIEW

[论文解读] Asymptotics of cut distributions and robust modular inference using Posterior Bootstrap

Emilia Pompe, Kasprzak, Mikołaj J.|ArXiv.org|Oct 21, 2021
Statistical Methods and Bayesian Inference参考文献 49被引用 6
一句话总结

本文为模块化贝叶斯推断中的截断分布建立了渐近理论,表明标准可信区域缺乏正确的频率学覆盖。提出了一种基于后部自展法的算法,通过可并行化的数值优化,生成具有名义频率学覆盖的可信区间,从而纠正这一问题。

ABSTRACT

Bayesian inference provides a framework to combine various model components with shared parameters, allowing joint uncertainty estimation and the use of all available data sources. Unfortunately, misspecification of any part of the model might propagate to all other parts and can lead to unsatisfactory results. Cut distributions have been proposed as a remedy, where the information is prevented from flowing along certain directions. We study cut distributions from an asymptotic perspective and obtain a Bernstein-von Mises theorem, as well as a Laplace approximation with quantitative bounds. We then propose an algorithm based on the Posterior Bootstrap that delivers credible regions with the nominal frequentist asymptotic coverage. The proposed methods are illustrated with numerical experiments in a variety of examples, including causal inference with propensity scores.

研究动机与目标

  • 解决模块化贝叶斯模型中由截断后验分布推导出的可信区域缺乏频率学覆盖的问题。
  • 提供一种稳健、模块化的推断框架,防止错误指定模块的信息流向可信模块。
  • 开发一种计算高效的可信区域构造方法,确保具有正确的渐近频率学覆盖。
  • 通过高维或复杂模块化模型中的可并行化数值优化,实现可扩展推断。
  • 在因果推断、流行病学和药代动力学/药效学建模等模型误设常见的场景中,展示该方法的实用性。

提出的方法

  • 使用拉普拉斯近似推导截断后验的渐近展开,识别标准可信区域中的偏差。
  • 提出一种后部自展算法,通过对后验样本重新加权,校正截断分布中的覆盖偏差。
  • 利用数值优化计算后部自展权重,实现在各模块间完全可并行化的计算。
  • 将该方法应用于在模型误设下估计第二模块的预测分布,校正反馈偏差。
  • 在渐近框架下,建立后部自展法与真实截断后验之间的理论等价性。
  • 利用Kullback-Leibler散度展开,比较贝叶斯模块化推断与后部自展法的预测性能。
Figure 1 : Density contours, corresponding to the 95% mass region, for samples obtained using PBMI (Algorithm 1 and Algorithm 2 ) and Bayesian modular inference for model ( 18 ). We considered three settings, depending on the choice of $\rho$ and $\sigma^{2}$ in the data-generating mechanism. Each p
Figure 1 : Density contours, corresponding to the 95% mass region, for samples obtained using PBMI (Algorithm 1 and Algorithm 2 ) and Bayesian modular inference for model ( 18 ). We considered three settings, depending on the choice of $\rho$ and $\sigma^{2}$ in the data-generating mechanism. Each p

实验结果

研究问题

  • RQ1为何标准截断分布无法在模块化贝叶斯推断中实现正确的频率学覆盖?
  • RQ2基于后部自展法的方法能否校正截断分布可信区域的覆盖不足?
  • RQ3在预测准确性方面,后部自展法与标准贝叶斯模块化推断相比表现如何?
  • RQ4在模型误设下,后部自展法的渐近行为如何?
  • RQ5所提出的方法能否在大规模模块化模型中高效实现并行化?

主要发现

  • 由于拉普拉斯近似中的偏差,标准截断分布产生的可信区域具有错误的频率学覆盖。
  • 后部自展法通过校正后验方差结构,生成具有渐近正确频率学覆盖的可信区域。
  • 该方法涉及对每个模块独立求解优化问题,从而实现完全并行化。
  • 在 $\sigma^2 = 1$ 的误设正态模型中,后部自展法将KL散度中的迹项从 -0.25 减少至 -0.21,改善了预测性能。
  • 当 $\sigma^2 = 2$ 时,后部自展法的迹值达到 -0.36,优于标准贝叶斯方法的预测准确性。
  • 理论展开表明,后部自展法与真实截断后验的渐近行为在 $o(n^{-1})$ 项内一致。
Figure 2 : Boxplots of elpd values obtained for model ( 19 ). Each plot is based on 50 datasets generated with $\sigma_{1}^{2}=1$ , $\sigma_{2}^{2}=0.5$ , $n_{2}$ presented on the $x$ -axis and $n_{1}=n_{2}/10$ . We use $N=2000$ samples drawn using each method.
Figure 2 : Boxplots of elpd values obtained for model ( 19 ). Each plot is based on 50 datasets generated with $\sigma_{1}^{2}=1$ , $\sigma_{2}^{2}=0.5$ , $n_{2}$ presented on the $x$ -axis and $n_{1}=n_{2}/10$ . We use $N=2000$ samples drawn using each method.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。