Skip to main content
QUICK REVIEW

[论文解读] Uncertainty in Bayesian Leave-One-Out Cross-Validation Based Model Comparison

Tuomas Sivula, Måns Magnusson|arXiv (Cornell University)|Aug 24, 2020
Statistical Methods and Bayesian Inference参考文献 24被引用 100
一句话总结

这篇论文分析用于比较两个模型的贝叶斯 LOO-CV 的不确定性,表明在有限样本下标准误估计可能不可靠,尤其当模型相似、错误指定或数据稀缺时,并提出具有实际指导意义的正态自助法和贝叶斯自助法方法。

ABSTRACT

It is useful to estimate the expected predictive performance of models planned to be used for prediction. We focus on leave-one-out cross-validation (LOO-CV), which has become a popular method for estimating predictive performance of Bayesian models. Given two models, we are interested in comparing the predictive performances and associated uncertainty, which can also be used to compute the probability of one model having better predictive performance than the other model. We study the properties of the Bayesian LOO-CV estimator and the related uncertainty quantification for the predictive performance difference, and analyse when a normal approximation of this uncertainty is well calibrated and whether taking into account higher moments could improve the approximation. We provide new results of the properties both theoretically in the linear regression case and empirically for hierarchical linear, latent linear, and spline models and discuss the challenges. We show that problematic cases include: comparing models with similar predictions, misspecified models, and small data. In these cases, there is a weak connection between the distributions of the LOO-CV estimator and its error. We show that that the problematic skewness of the error distribution for the difference, which occurs when the models make similar predictions, does not fade away when the data size grows to infinity in certain situations. Based on the results, we also provide some practical recommendations for the users of Bayesian LOO-CV for comparing predictive performance of models.

研究动机与目标

  • 评估在使用贝叶斯 LOO-CV 进行模型比较时,elpd 差异的不确定性如何表现。
  • 识别标准不确定性估计在何种情形下不可靠(例如预测相似、模型不正确指定、数据量小)。
  • 分析在正态线性回归及其他模型中,LOO-CV 不确定性的理论与经验特性。
  • 为使用贝叶斯 LOO-CV 的实践者提供实用建议。

提出的方法

  • 为模型比较形式化 elpd 及其 LOO-CV 估计量。
  • 通过误差 err_LOO 及其分布 p(err_LOO) 分析差分 elpd(Ma, Mb|y) 的不确定性。
  • 比较两种近似方法:误差分布的正态近似以及贝叶斯自助法(Dirichlet)近似。
  • 推导正态线性回归的解析结果,并在多种模型中通过实验进行验证。
  • 使用 PIT(概率积分变换)评估近似不确定性相对于真分布(oracle 分布)的校准性。
  • 讨论渐近行为和有限样本问题,包括偏斜和错误指定。

实验结果

研究问题

  • RQ1在使用贝叶斯 LOO-CV 比较两个模型时,预测性能差异的标准不确定性估计有多可靠?
  • RQ2在哪些情形下正态近似或贝叶斯自助法近似会失效或校准性差?
  • RQ3偏斜、错误指定和小样本如何影响 LOO-CV 模型比较的不确定性?
  • RQ4结果是否可以推广到超越正态线性回归的其他模型,如分层模型、泊松广义线性模型和样条?

主要发现

  • 在有限样本中,LOO-CV 差异的不确定性估计可能很差,尤其当模型预测相近、存在错误指定或数据受限时。
  • LOO-CV 估计误差的分布可能高度偏斜,在某些情景下使得正态近似不可靠。
  • 错误指定和离群值可能偏置 LOO-CV 估计并使方差膨胀,影响模型比较的结论。
  • 即使数据量增大,某些有问题的偏斜模式也可能持续存在,妨碍对哪个模型更优的准确推断。
  • 贝叶斯自助法并不普遍优于正态近似来估计 elpd 差异的不确定性。
  • 正态线性回归的结果在定性上可推广到其他模型,且在贝叶斯 K 折交叉验证中观察到类似行为。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。