[论文解读] Robust Neural Posterior Estimation and Statistical Model Criticism
本文提出稳健神经后验估计(RNPE),通过显式建模模拟器输出与真实数据之间的差异,增强神经后验估计(NPE)方法,从而在模型设定错误时实现稳健推理与可解释的模型批评。在模拟和现实世界示例中,通过控制模型误差,RNPE 在模拟器设定错误时仍能保持准确且保守的后验分布,优于标准 NPE 方法。
Computer simulations have proven a valuable tool for understanding complex phenomena across the sciences. However, the utility of simulators for modelling and forecasting purposes is often restricted by low data quality, as well as practical limits to model fidelity. In order to circumvent these difficulties, we argue that modellers must treat simulators as idealistic representations of the true data generating process, and consequently should thoughtfully consider the risk of model misspecification. In this work we revisit neural posterior estimation (NPE), a class of algorithms that enable black-box parameter inference in simulation models, and consider the implication of a simulation-to-reality gap. While recent works have demonstrated reliable performance of these methods, the analyses have been performed using synthetic data generated by the simulator model itself, and have therefore only addressed the well-specified case. In this paper, we find that the presence of misspecification, in contrast, leads to unreliable inference when NPE is used naively. As a remedy we argue that principled scientific inquiry with simulators should incorporate a model criticism component, to facilitate interpretable identification of misspecification and a robust inference component, to fit 'wrong but useful' models. We propose robust neural posterior estimation (RNPE), an extension of NPE to simultaneously achieve both these aims, through explicitly modelling the discrepancies between simulations and the observed data. We assess the approach on a range of artificially misspecified examples, and find RNPE performs well across the tasks, whereas naively using NPE leads to misleading and erratic posteriors.
研究动机与目标
- 解决标准神经后验估计(NPE)在模拟器设定错误时产生不可靠推断的关键局限性。
- 开发一个统一框架,同时实现基于模拟器模型的稳健参数推断与可解释的模型批评。
- 提供诊断工具,用于识别哪些摘要统计量最可能设定错误,利用显式的设定错误概率估计。
- 将模型批评与推断解耦,避免将糟糕的推断与实际的模型设定错误混淆。
- 在模拟器无法完全捕捉真实数据生成过程的现实场景中,展示 RNPE 的实际效用。
提出的方法
- RNPE 通过引入一个稀疏-滑动误差模型,扩展了 NPE,显式表示每个摘要统计量设定错误的概率。
- 将观测数据建模为模拟器输出的扰动版本,利用去噪机制估计潜在的‘真实’数据分布。
- 通过分层贝叶斯模型,联合学习参数的后验分布与每个摘要统计量的设定错误状态。
- 使用 MCMC 从参数与设定错误指示变量的联合后验中抽样,即使模拟器错误也能实现稳健推断。
- 该框架支持诊断性解释:设定错误概率较高的摘要统计量被标记为模型差异的来源。
- 利用归一化流进行密度估计,实现高维设置下灵活且可扩展的后验近似。
实验结果
研究问题
- RQ1当观测数据远离模拟器支持范围时,标准神经后验估计(NPE)在模型设定错误下的表现如何,特别是当观测数据位于模拟器支持范围之外?
- RQ2能否开发一个统一框架,同时在基于模拟的推断中实现稳健参数推断与模型批评?
- RQ3显式设定错误概率诊断在多大程度上能帮助识别哪些摘要统计量最可能导致模型与数据的不一致?
- RQ4在人工设定错误的模拟器下,RNPE 与标准 NPE 在后验准确性与覆盖度方面的表现相比如何?
- RQ5所提出的方法能否在现实世界的模拟任务中检测并诊断模型设定错误,例如涉及生物或流行病学模型的任务?
主要发现
- 在高斯示例中,当观测数据超出模拟器支持范围时,标准 NPE 产生了过度自信且不稳定的后验分布,而 RNPE 保持了准确且保守的后验估计。
- 在具有报告延迟的 SIR 模型中,NPE 将后验集中在错误的参数值上,而 RNPE 正确识别了设定错误,并生成了更准确的后验分布。
- 在包含坏死肿瘤区域的 CS(癌症模拟)模型中,RNPE 正确将 N Cancer 摘要统计量标记为设定错误,与关于组织坏死的领域知识一致。
- RNPE 在所有任务中均表现出一致的性能,将更多后验质量集中在真实参数上,并在设定错误时展现出优于 NPE 的覆盖特性。
- 该方法成功识别了设定错误之间的权衡,例如当两个摘要统计量中仅一个可被良好设定时,揭示了模型拟合的不确定性。
- 稀疏-滑动误差模型提供了可解释的诊断,设定错误概率清晰地指明了哪些统计量与观测数据最不一致。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。