Skip to main content
QUICK REVIEW

[Paper Review] Leave-One-Out Cross-Validation for Bayesian Model Comparison in Large Data

Måns Magnusson, Michael Riis Andersen|arXiv (Cornell University)|Jan 3, 2020
Statistical Methods and Bayesian Inference32 references30 citations
TL;DR

The paper develops an efficient approach for comparing Bayesian models on large datasets by combining a difference estimator with fast LOO surrogates, enabling accurate elpd differences with minimal subsampling and model-agnostic sampling.

ABSTRACT

Recently, new methods for model assessment, based on subsampling and posterior approximations, have been proposed for scaling leave-one-out cross-validation (LOO) to large datasets. Although these methods work well for estimating predictive performance for individual models, they are less powerful in model comparison. We propose an efficient method for estimating differences in predictive performance by combining fast approximate LOO surrogates with exact LOO subsampling using the difference estimator and supply proofs with regards to scaling characteristics. The resulting approach can be orders of magnitude more efficient than previous approaches, as well as being better suited to model comparison.

Motivation & Objective

  • Motivate and quantify the need for scalable Bayesian model comparison using elpd (expected log predictive density).
  • Develop an efficient procedure to estimate differences in elpd between models with large data via a difference estimator and subsampling.
  • Incorporate effective parameter count p_eff into LOO surrogates to improve accuracy.
  • Provide theoretical guarantees for convergence and discuss computational trade-offs for large-data scenarios.

Proposed method

  • Use the difference estimator with simple random sampling without replacement to estimate elpd differences between models (Eq. 7).
  • Augment LOO surrogates with p_eff to improve approximation quality (Eq. 1.2 and related discussion).
  • Propose fast approximate surrogates (Delta WAIC variants, Taylor-based p_eff approximations, PSIS-LOO alternatives) to obtain pi_tilde with reduced cost (Sec. 2.2).
  • Demonstrate that the difference estimator yields unbiased estimates of elpd_loo and its variance and that convergence holds as pi_tilde converges in mean to pi (Propositions 2 and 3).
  • Outline computational cost for various surrogates and show how costs scale with n, P, and S (Table 1).

Experimental results

Research questions

  • RQ1Does using better approximations pi_tilde improve empirical performance in elpd estimation and model comparison?
  • RQ2Which pi_tilde surrogate provides favorable cost–accuracy trade-offs for large data?
  • RQ3How does the difference estimator compare with the Hansen-Hurwitz (HH) approach for estimating elpd differences and variances?
  • RQ4How well does the method scale when performing large-data Bayesian model comparisons?

Key findings

  • Including p_eff in the estimation of elpd markedly improves accuracy (often by orders of magnitude) compared to using plpd alone.
  • The difference estimator enables reusing a single subsample to estimate multiple model comparisons, reducing computational cost.
  • HH and the difference estimator perform similarly for individual elpd estimates, though HH can marginally outperform for some models; the advantage of the difference estimator lies in estimating variances and model comparisons.
  • TIS-based surrogates (e.g., TIS_2k) provide the best accuracy among surrogates for large or hierarchical models, at a manageable cost; plpd is recommended for simpler models.
  • The method demonstrates favorable scaling where a small subsample (e.g., m around 100–400) suffices to compare models, with accuracy improving as pi_tilde and p_eff approximations improve.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.