Skip to main content
QUICK REVIEW

[Paper Review] Limitations of "Limitations of Bayesian leave-one-out cross-validation for model selection"

Aki Vehtari, Daniel Simpson|arXiv (Cornell University)|Oct 12, 2018
Statistical Methods and Bayesian InferenceMathematics17 references21 citations
TL;DR

This paper challenges the criticisms of Bayesian leave-one-out cross-validation (LOO-CV) in Gronau and Wagenmakers (2018), arguing that their interpretation misrepresents the standard, recommended use of LOO. It advocates for using LOO with uncertainty estimates and model weights via pseudo-BMA+ or stacking, emphasizing that LOO should not be reduced to a single-number decision rule but rather used with diagnostics like the $ˆ{k}$ statistic to assess model fit and predictive performance.

ABSTRACT

This article is an invited discussion of the article by Gronau and Wagenmakers (2018) that can be found at https://dx.doi.org/10.1007/s42113-018-0011-7.

Motivation & Objective

  • To correct misrepresentations of Bayesian leave-one-out cross-validation (LOO-CV) in Gronau and Wagenmakers (2018), particularly the claim that LOO is fundamentally flawed.
  • To argue that LOO should not be used as a single-number decision rule but rather with uncertainty quantification and diagnostic checks.
  • To promote the use of pseudo-BMA+ weights and stacking as superior alternatives to pseudo-Bayes factors for model weighting in LOO.
  • To highlight the importance of diagnostics such as the $ˆ{k}$ statistic for detecting influential observations and model misspecification.
  • To emphasize that LOO is better suited for model criticism than for forced model selection, especially when no model is clearly correct.

Proposed method

  • Reinterprets LOO-CV as an estimate of expected log predictive density, with uncertainty quantified via empirical variance of the LOO estimate.
  • Recommends replacing pseudo-Bayes factors with pseudo-BMA+ weights, which assume normality of the log predictive density differences and adjust model weights accordingly.
  • Proposes stacking weights as a more principled method for model combination that directly applies the LOO principle.
  • Uses decomposition of LOO differences into individual data-point contributions to identify problematic observations and model-data misfit.
  • Applies the $ˆ{k}$ diagnostic to detect high leverage points where the predictive distribution changes drastically when an observation is excluded.
  • Stresses the importance of conditional exchangeability in LOO and warns against using LOO in time-series, spatial, or hierarchical models without block-based cross-validation.

Experimental results

Research questions

  • RQ1Why is the use of a single-number summary, such as a pseudo-Bayes factor, problematic for LOO-based model selection?
  • RQ2How can uncertainty in LOO estimates be properly accounted for in model comparison?
  • RQ3What are the advantages of pseudo-BMA+ and stacking weights over pseudo-Bayes factors in LOO-based model averaging?
  • RQ4In what ways can LOO diagnostics like $ˆ{k}$ improve model criticism and reveal model misspecification?
  • RQ5Under what conditions does LOO fail, and how can these failures be detected or avoided?

Key findings

  • The pseudo-Bayes factor used in Gronau and Wagenmakers (2018) is not representative of standard LOO practice, which should include uncertainty quantification and model weights.
  • Pseudo-BMA+ weights, which assume normality of the log predictive density differences, are a more appropriate and robust alternative to pseudo-Bayes factors.
  • Stacking weights provide a more principled method for model combination by directly applying the LOO principle, improving predictive performance.
  • The $ˆ{k}$ diagnostic effectively identifies high-leverage observations where model predictions change significantly when the data point is omitted.
  • LOO is not inherently flawed but fails when the data is not exchangeable conditional on model parameters or when future data differs substantially from observed data.
  • LOO can express epistemological uncertainty by refusing to select a single model when all models are inadequate, unlike marginal likelihood methods that always favor one model regardless of fit.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.