Skip to main content
QUICK REVIEW

[论文解读] Simulation-Based Calibration Checking for Bayesian Computation: The Choice of Test Quantities Shapes Sensitivity

Martin Modrák, Angie H. Moon|arXiv (Cornell University)|Nov 4, 2022
Statistical Methods and Bayesian Inference被引用 12
一句话总结

本文提出了一种新型基于模拟的校准(SBC)方法,通过引入与数据相关的检验统计量,特别是数据与参数的联合似然,增强了贝叶斯推断中对计算误差的敏感度。该方法通过重新定义SBC以使用灵活的检验统计量,能够检测此前无法察觉的问题(如后验分布与先验分布相等),显著提升了验证能力,同时提供了理论基础和SBC R包中的实际实现。

ABSTRACT

Simulation-based calibration checking (SBC) is a practical method to validate computationally-derived posterior distributions or their approximations. In this paper, we introduce a new variant of SBC to alleviate several known problems. Our variant allows the user to in principle detect any possible issue with the posterior, while previously reported implementations could never detect large classes of problems including when the posterior is equal to the prior. This is made possible by including additional data-dependent test quantities when running SBC. We argue and demonstrate that the joint likelihood of the data is an especially useful test quantity. Some other types of test quantities and their theoretical and practical benefits are also investigated. We provide theoretical analysis of SBC, thereby providing a more complete understanding of the underlying statistical mechanisms. We also bring attention to a relatively common mistake in the literature and clarify the difference between SBC and checks based on the data-averaged posterior. We support our recommendations with numerical case studies on a multivariate normal example and a case study in implementing an ordered simplex data type for use with Hamiltonian Monte Carlo. The SBC variant introduced in this paper is implemented in the $\mathtt{SBC}$ R package.

研究动机与目标

  • 为解决标准SBC在无法检测到重大计算错误(例如后验分布等于先验分布)方面的局限性。
  • 通过引入与数据相关的检验统计量(尤其是数据与参数的联合似然),提升SBC的敏感度。
  • 为理解检验统计量选择如何影响SBC性能,提供一个理论严谨的框架。
  • 澄清文献中常见的误解,特别是区分SBC与数据平均后验检查。
  • 通过数值案例研究和集成到SBC R包中,支持实际应用。

提出的方法

  • 提出SBC的一种变体,除参数本身外,还包含额外的与数据相关的检验统计量。
  • 由于对模型和计算错误具有强敏感性,使用联合似然π_joint(y,θ)作为特别有效的检验统计量。
  • 推导出检验统计量满足SBC的理论条件,包括分位数函数和累积分布函数的约束。
  • 提出一个分析检验统计量族及其对SBC敏感度影响的框架,使用随机序关系和函数约束。
  • 通过多元正态模型和哈密顿蒙特卡洛实现的数值案例研究,验证该方法。
  • 在SBC R包中实现新的SBC变体,以支持实际应用中的贝叶斯计算验证。

实验结果

研究问题

  • RQ1检验统计量的选择如何影响基于模拟的校准对贝叶斯推断中计算误差的敏感度?
  • RQ2SBC能否检测到后验分布与先验分布完全相同的情况——这一故障模式此前未被标准SBC检测到?
  • RQ3检验统计量需满足何种理论条件,才能确保在不同模型结构下SBC性能的有效性?
  • RQ4与其它检验统计量相比,数据与参数的联合似然在误差检测能力方面表现如何?
  • RQ5在真实世界的贝叶斯计算中,使用与数据相关的检验统计量对SBC有何实际影响?

主要发现

  • 所提出的SBC变体能够检测后验分布中的任何计算错误,包括后验分布与先验分布完全相同的情况——此前标准SBC无法检测到。
  • 使用联合似然作为检验统计量可显著提高敏感度,因为它对所有数据取值下的后验分布施加了强约束。
  • 理论分析表明,检验统计量的选择从根本上决定了SBC检测偏差的能力,某些选择会施加后验族的严格函数约束。
  • 在离散参数空间中,使用参数本身等检验统计量的SBC会导致解空间高度受限,通常迫使后验分布与真实后验或先验相匹配。
  • 数值案例研究证实,新方法能检测到标准SBC遗漏的错误,特别是在复杂模型(如HMC的有序单纯形实现)中。
  • 在SBC R包中的实现使研究人员能够轻松将增强后的SBC方法应用于实际后验近似验证。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。