Skip to main content
QUICK REVIEW

[论文解读] Diagnosing model misspecification and performing generalized Bayes' updates via probabilistic classifiers

Owen Thomas, Jukka Corander|arXiv (Cornell University)|Dec 12, 2019
Gaussian Processes and Bayesian Inference参考文献 16被引用 8
一句话总结

该论文提出使用概率分类器来诊断贝叶斯推断中的模型误设问题,并通过似然退火引导广义贝叶斯更新。通过训练分类器以区分基于统计模型模拟的数据与实际观测数据,该方法估计对数似然比,其期望值近似于模型与真实数据生成过程之间的负Kullback-Leibler散度,从而实现稳健的模型评估和最优退火选择,而无需依赖真实模型。

ABSTRACT

Model misspecification is a long-standing enigma of the Bayesian inference framework as posteriors tend to get overly concentrated on ill-informed parameter values towards the large sample limit. Tempering of the likelihood has been established as a safer way to do updates from prior to posterior in the presence of model misspecification. At one extreme tempering can ignore the data altogether and at the other extreme it provides the standard Bayes' update when no misspecification is assumed to be present. However, it is an open issue how to best recognize misspecification and choose a suitable level of tempering without access to the true generating model. Here we show how probabilistic classifiers can be employed to resolve this issue. By training a probabilistic classifier to discriminate between simulated and observed data provides an estimate of the ratio between the model likelihood and the likelihood of the data under the unobserved true generative process, within the discriminatory abilities of the classifier. The expectation of the logarithm of a ratio with respect to the data generating process gives an estimation of the negative Kullback-Leibler divergence between the statistical generative model and the true generative distribution. Using a set of canonical examples we show that this divergence provides a useful misspecification diagnostic, a model comparison tool, and a method to inform a generalised Bayesian update in the presence of misspecification for likelihood-based models.

研究动机与目标

  • 解决贝叶斯推断中模型误设的挑战,即在大样本极限下后验分布集中在参数值较差的情况。
  • 在真实数据生成模型未知的情况下,开发一种选择最优退火水平t的广义贝叶斯更新方法。
  • 提供一种无需访问真实模型或显式计算似然的模型设定诊断工具。
  • 将基于分类器的密度比估计整合到贝叶斯推断中,以提升在模型误设下的鲁棒性。

提出的方法

  • 训练一个概率分类器,以区分在不同参数值下从统计模型模拟的数据与实际观测数据。
  • 利用分类器的几率比估计模型似然与真实数据生成似然的比值,从而实现密度比估计。
  • 计算在真实数据生成过程中对数几率比的期望值,作为模型与真实分布之间负Kullback-Leibler散度的近似。
  • 将估计的KL散度用作模型误设的诊断工具,值越大表示偏差越大。
  • 对估计的对数比值进行单尾t检验,以评估模型是否显著误设(p值 < 0.05 表示存在误设)。
  • 利用该诊断结果指导广义后验更新中退火参数t的选择,避免在误设模型中产生过度自信。

实验结果

研究问题

  • RQ1在缺乏真实数据生成模型的情况下,概率分类器能否可靠检测模型误设?
  • RQ2基于分类器的对数似然比估计在多大程度上能近似真实模型与真实分布之间的负Kullback-Leibler散度?
  • RQ3估计的散度能否用于指导广义贝叶斯更新中最优退火参数t的选择?
  • RQ4当真实模型位于模型类之外时,该方法对分类器校准和模型复杂度的鲁棒性如何?

主要发现

  • 基于分类器的对数几率比期望估计值与统计模型和真实数据生成过程之间的负Kullback-Leibler散度高度近似,提供了有效的误设诊断。
  • 对估计对数比值进行单尾t检验可提供正式的假设检验,p值低于0.05表示模型拟合显著不佳。
  • 在所有测试的典型模型(正态分布、泊松分布和线性回归)中,当真实分布与假设模型不同时(例如拉普拉斯分布 vs. 正态分布,负二项分布 vs. 泊松分布),该方法成功检测到误设。
  • 该方法能够识别出平衡先验信息与数据证据的最优退火水平t,从而减少在误设模型中的过度自信。
  • 即使在人类直觉可能失效的复杂判别任务中,该方法仍保持良好条件性,且在假设检验中KL散度的低估是保守的。
  • 补充图表明,在多种模型类型和退火水平下,真实值与分类器估计的对数比值之间具有高度一致性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。