Skip to main content
QUICK REVIEW

[Paper Review] A Bayes interpretation of stacking for M-complete and M-open settings

Tri Le, Bertrand Clarke|arXiv (Cornell University)|Feb 16, 2016
Bayesian Modeling and Causal Inference11 references3 citations
TL;DR

This paper provides a Bayesian justification for stacking in both M-complete and M-open settings by showing that stacking weights asymptotically minimize posterior expected loss, thereby formalizing cross-validation as a Bayes-optimal procedure. It relaxes standard positivity and sum-to-one constraints on weights and proposes data-driven basis generation via bootstrap sampling to optimize predictor selection and weighting.

ABSTRACT

In M-open problems where no true model can be conceptualized, it is common to back off from modeling and merely seek good prediction. Even in M-complete problems, taking a predictive approach can be very useful. Stacking is a model averaging procedure that gives a composite predictor by combining individual predictors from a list of models using weights that optimize a cross-validation criterion. We show that the stacking weights also asymptotically minimize a posterior expected loss. Hence we formally provide a Bayesian justification for cross-validation. Often the weights are constrained to be positive and sum to one. For greater generality, we omit the positivity constraint and relax the `sum to one' constraint. A key question is `What predictors should be in the average?' We first verify that the stacking error depends only on the span of the models. Then we propose using bootstrap samples from the data to generate empirical basis elements that can be used to form models. We use this in two computed examples to give stacking predictors that are (i) data driven, (ii) optimal with respect to the number of component predictors, and (iii) optimal with respect to the weight each predictor gets.

Motivation & Objective

  • To formalize stacking as a Bayes-optimal procedure in M-complete and M-open statistical settings.
  • To relax conventional constraints on stacking weights (positivity and sum-to-one) for greater generality.
  • To develop a data-driven method for selecting and weighting predictors using bootstrap-generated empirical basis elements.
  • To demonstrate that stacking error depends only on the span of the models, enabling model-agnostic optimization.
  • To provide theoretical justification for using cross-validation as a proxy for posterior expected loss minimization in predictive modeling.

Proposed method

  • Derives the asymptotic equivalence between minimizing posterior expected loss and optimizing stacking weights via leave-one-out cross-validation.
  • Relaxes standard constraints on stacking weights by removing positivity and sum-to-one requirements, allowing for more flexible optimization.
  • Uses bootstrap resampling to generate empirical basis elements that represent the span of candidate models, enabling data-driven model construction.
  • Applies Theorem 3.7 to show that stacking error depends only on the model span, not on the specific basis used.
  • Employs orthonormal decomposition and projection theory in Hilbert space to prove that minimum stacking error is invariant across equivalent model bases.
  • Uses the relationship between leverage, residuals, and predicted values to derive closed-form expressions for stacking weights under specific assumptions.

Experimental results

Research questions

  • RQ1Can stacking be formally justified as a Bayes-optimal procedure in M-complete and M-open settings?
  • RQ2How do relaxed constraints on stacking weights (no positivity or sum-to-one) affect predictive performance and theoretical properties?
  • RQ3What is the role of model span in determining stacking error, and how can it be leveraged for optimal predictor construction?
  • RQ4Can bootstrap-based empirical basis elements improve the selection and weighting of component predictors in stacking?
  • RQ5To what extent does leave-one-out cross-validation approximate posterior expected loss minimization in predictive modeling?

Key findings

  • Stacking weights asymptotically minimize posterior expected loss, providing a formal Bayesian justification for cross-validation in M-complete problems.
  • The stacking error depends only on the span of the models, not on the specific basis used, implying that model selection can be decoupled from basis representation.
  • When leverage values vanish asymptotically, stacking weights converge to equal weights (1/2 each) under symmetric residual and prediction structures.
  • With a sum-to-two constraint, equal prediction variances across models lead to equal weights of 1, demonstrating adaptability to constraint changes.
  • The minimum stacking error is invariant across different orthonormal bases of the same model span, proving theoretical robustness.
  • Theoretical results extend to leave-k-out cross-validation, supporting broader applicability of the Bayesian justification for cross-validation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.