[Paper Review] Uncertainty in Bayesian Leave-One-Out Cross-Validation Based Model Comparison
The paper analyzes the uncertainty of Bayesian LOO-CV used for comparing two models, showing that standard error estimates can be unreliable in finite samples, especially when models are similar, misspecified, or data are scarce, and proposes normal and Bayesian bootstrap approaches with practical guidance.
It is useful to estimate the expected predictive performance of models planned to be used for prediction. We focus on leave-one-out cross-validation (LOO-CV), which has become a popular method for estimating predictive performance of Bayesian models. Given two models, we are interested in comparing the predictive performances and associated uncertainty, which can also be used to compute the probability of one model having better predictive performance than the other model. We study the properties of the Bayesian LOO-CV estimator and the related uncertainty quantification for the predictive performance difference, and analyse when a normal approximation of this uncertainty is well calibrated and whether taking into account higher moments could improve the approximation. We provide new results of the properties both theoretically in the linear regression case and empirically for hierarchical linear, latent linear, and spline models and discuss the challenges. We show that problematic cases include: comparing models with similar predictions, misspecified models, and small data. In these cases, there is a weak connection between the distributions of the LOO-CV estimator and its error. We show that that the problematic skewness of the error distribution for the difference, which occurs when the models make similar predictions, does not fade away when the data size grows to infinity in certain situations. Based on the results, we also provide some practical recommendations for the users of Bayesian LOO-CV for comparing predictive performance of models.
Motivation & Objective
- Assess how uncertainty in elpd differences behaves when using Bayesian LOO-CV for model comparison.
- Identify situations where standard uncertainty estimates are unreliable (e.g., similar predictions, misspecification, small data).
- Analyze theoretical and empirical properties of LOO-CV uncertainty in normal linear regression and other models.
- Provide practical recommendations for practitioners using Bayesian LOO-CV.
Proposed method
- Formulate elpd and its LOO-CV estimator for model comparison.
- Analyze uncertainty in the difference elpd(Ma, Mb|y) via error err_LOO and its distribution p(err_LOO).
- Compare two approximation approaches: normal approximation and Bayesian bootstrap (Dirichlet) for the error distribution.
- Derive analytical results for normal linear regression and validate with experiments across multiple models.
- Use PIT to assess calibration of the approximated uncertainty against the oracle distribution.
- Discuss asymptotic behavior and finite-sample issues including skewness and misspecification.
Experimental results
Research questions
- RQ1How reliable are standard uncertainty estimates for the difference in predictive performance when using Bayesian LOO-CV to compare two models?
- RQ2In which scenarios do normal or Bayesian bootstrap approximations fail or become poorly calibrated?
- RQ3How do skewness, misspecification, and small sample size affect the uncertainty in LOO-CV model comparison?
- RQ4Do the results generalize beyond normal linear regression to other models like hierarchical, Poisson GLM, and splines?
Key findings
- Uncertainty in LOO-CV differences can be poorly estimated in finite samples, especially when models have similar predictions, are misspecified, or data are limited.
- The distribution of the LOO-CV estimator error can be highly skewed, making normal approximations unreliable in some scenarios.
- Misspecification and outliers can bias the LOO-CV estimate and inflate variance, affecting model comparison conclusions.
- Even with larger data sizes, certain problematic skewness patterns may persist, hindering accurate inference about which model is better.
- Bayesian bootstrap does not universally outperform the normal approximation in practice for the uncertainty of elpd differences.
- The results for normal linear regression extend qualitatively to other models, and similar behavior is observed in Bayesian K-fold CV.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.