Skip to main content
QUICK REVIEW

[Paper Review] Aggregating distribution forecasts from deep ensembles

Benedikt Schulz, Köhler, Lutz|arXiv (Cornell University)|Apr 5, 2022
Wind and Air Flow Studies6 citations
TL;DR

This paper proposes a quantile-based aggregation framework for deep ensembles in probabilistic forecasting, demonstrating that quantile aggregation (Vincentization) significantly outperforms linear pooling of density forecasts. Using theoretical analysis, simulations, and a wind gust forecasting case study, the authors show that aggregating forecast distributions via quantiles improves predictive performance, especially when ensemble sizes are moderate (10–20 members) and individual forecasts are well-optimized.

ABSTRACT

The importance of accurately quantifying forecast uncertainty has motivated much recent research on probabilistic forecasting. In particular, a variety of deep learning approaches has been proposed, with forecast distributions obtained as output of neural networks. These neural network-based methods are often used in the form of an ensemble, e.g., based on multiple model runs from different random initializations or more sophisticated ensembling strategies such as dropout, resulting in a collection of forecast distributions that need to be aggregated into a final probabilistic prediction. With the aim of consolidating findings from the machine learning literature on ensemble methods and the statistical literature on forecast combination, we address the question of how to aggregate distribution forecasts based on such `deep ensembles'. Using theoretical arguments and a comprehensive analysis on twelve benchmark data sets, we systematically compare probability- and quantile-based aggregation methods for three neural network-based approaches with different forecast distribution types as output. Our results show that combining forecast distributions from deep ensembles can substantially improve the predictive performance. We propose a general quantile aggregation framework for deep ensembles that allows for corrections of systematic deficiencies and performs well in a variety of settings, often superior compared to a linear combination of the forecast densities. Finally, we investigate the effects of the ensemble size and derive recommendations of aggregating distribution forecasts from deep ensembles in practice.

Motivation & Objective

  • To consolidate statistical and machine learning literature on forecast combination and ensembling for probabilistic forecasting.
  • To systematically compare probability- and quantile-based aggregation methods for deep ensembles with different neural network-based forecast distribution types.
  • To evaluate the impact of ensemble size on predictive performance and derive practical recommendations for aggregating distribution forecasts.
  • To investigate whether quantile aggregation (Vincentization) provides superior performance compared to linear pooling of densities in deep ensemble settings.

Proposed method

  • The authors use a two-step workflow: first generating an ensemble of probabilistic forecasts via multiple random initializations of neural networks, then aggregating them into a single forecast distribution.
  • They compare two main aggregation strategies: linear pooling (LP) of forecast densities and Vincentization (VI), which linearly combines quantile functions across ensemble members.
  • The framework is applied to three distinct neural network architectures that output different types of forecast distributions: parametric (e.g., Gaussian), semi-parametric (quantile function approximations), and nonparametric (histogram-based).
  • Theoretical analysis is used to justify the shape-preservation property of Vincentization, particularly when ensemble members stem from the same model and data.
  • Simulation experiments and a real-world case study on probabilistic wind gust forecasting are conducted to evaluate and compare the performance of aggregation methods.
  • Performance is assessed using proper scoring rules (e.g., CRPS), and the impact of ensemble size on predictive accuracy is systematically evaluated.

Experimental results

Research questions

  • RQ1How do different aggregation methods—linear pooling of densities versus quantile-based Vincentization—affect the predictive performance of deep ensemble forecasts?
  • RQ2What is the optimal ensemble size for deep ensembles in probabilistic forecasting, balancing performance gains and computational cost?
  • RQ3Does the choice of neural network architecture (parametric, semi-parametric, nonparametric) influence the effectiveness of aggregation methods?
  • RQ4Under what conditions does quantile aggregation (VI) outperform linear pooling (LP) in terms of forecast calibration and sharpness?
  • RQ5Can forecast combination correct for systematic errors in individual ensemble members, or is it limited by their underlying model quality?

Key findings

  • Quantile-based aggregation via Vincentization (VI) consistently outperforms linear pooling of forecast densities in terms of predictive performance, especially in terms of CRPS and calibration.
  • Aggregating forecast distributions leads to substantial improvements in predictive accuracy, with the greatest gains observed when individual ensemble members are well-optimized.
  • An ensemble size of at least 10 members is recommended, as larger ensembles (beyond 20) yield only marginal performance gains, making 10–20 a computationally efficient sweet spot.
  • The superiority of VI over LP is most pronounced when ensemble members are based on the same model and data, where shape-preservation is a desirable property.
  • Forecast combination cannot fully correct for substantial systematic errors in individual forecasts, emphasizing the importance of optimizing base models before aggregation.
  • Dropout-based ensembles were found to yield worse performance than deep ensembles from random initialization, suggesting that the choice of ensemble generation method significantly impacts final forecast quality.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.