[论文解读] Aggregating distribution forecasts from deep ensembles
本文提出了一种基于分位数的深度集成概率预测聚合框架,表明分位数聚合(Vincentization)显著优于密度预测的线性池化方法。通过理论分析、模拟实验和风速突风预测案例研究,作者证明了基于分位数的聚合方法能提升预测性能,尤其在中等规模集成(10–20个成员)且各预测模型已充分优化时效果更佳。
The importance of accurately quantifying forecast uncertainty has motivated much recent research on probabilistic forecasting. In particular, a variety of deep learning approaches has been proposed, with forecast distributions obtained as output of neural networks. These neural network-based methods are often used in the form of an ensemble, e.g., based on multiple model runs from different random initializations or more sophisticated ensembling strategies such as dropout, resulting in a collection of forecast distributions that need to be aggregated into a final probabilistic prediction. With the aim of consolidating findings from the machine learning literature on ensemble methods and the statistical literature on forecast combination, we address the question of how to aggregate distribution forecasts based on such `deep ensembles'. Using theoretical arguments and a comprehensive analysis on twelve benchmark data sets, we systematically compare probability- and quantile-based aggregation methods for three neural network-based approaches with different forecast distribution types as output. Our results show that combining forecast distributions from deep ensembles can substantially improve the predictive performance. We propose a general quantile aggregation framework for deep ensembles that allows for corrections of systematic deficiencies and performs well in a variety of settings, often superior compared to a linear combination of the forecast densities. Finally, we investigate the effects of the ensemble size and derive recommendations of aggregating distribution forecasts from deep ensembles in practice.
研究动机与目标
- 整合统计学与机器学习领域关于概率预测中预报组合与集成方法的文献。
- 系统比较基于概率与分位数的聚合方法在不同基于神经网络的预报分布类型下的深度集成表现。
- 评估集成规模对预测性能的影响,并为预报分布的聚合提供实用建议。
- 探究在深度集成设置中,分位数聚合(Vincentization)是否优于密度线性池化方法。
提出的方法
- 作者采用两步工作流程:首先通过神经网络的多次随机初始化生成概率预测集合,然后将这些预测聚合为单一的预报分布。
- 比较两种主要聚合策略:预报密度的线性池化(LP)与基于分位数函数的线性组合(Vincentization,VI)。
- 该框架应用于三种不同的神经网络架构,分别输出不同类型的预报分布:参数型(如正态分布)、半参数型(分位数函数近似)以及非参数型(基于直方图)。
- 通过理论分析证明Vincentization具有形状保持特性,尤其当集成成员源自同一模型和数据时。
- 开展模拟实验和一个真实世界中的概率风速突风预测案例研究,以评估并比较不同聚合方法的性能。
- 使用合适的评分规则(如CRPS)评估性能,并系统分析集成规模对预测精度的影响。
实验结果
研究问题
- RQ1不同的聚合方法——密度线性池化与基于分位数的Vincentization——如何影响深度集成预测的预测性能?
- RQ2在性能提升与计算成本之间取得平衡时,深度集成在概率预测中的最优集成规模是多少?
- RQ3神经网络架构的选择(参数型、半参数型、非参数型)是否会影响聚合方法的有效性?
- RQ4在何种条件下,分位数聚合(VI)在预报校准性和精确性方面优于线性池化(LP)?
- RQ5预报组合能否纠正个体集成成员中的系统性误差,还是其表现受限于基础模型的质量?
主要发现
- 基于分位数的聚合方法(即Vincentization,VI)在预测性能方面始终优于密度线性池化方法,尤其在CRPS和校准性方面表现更优。
- 对预报分布进行聚合可显著提升预测精度,且当个体集成成员已充分优化时,性能提升最为明显。
- 建议集成规模至少为10个成员,因为当成员数超过20时,性能增益已趋于边际化,因此10–20个成员是计算效率的最优区间。
- 当集成成员基于同一模型和数据时,VI相对于LP的优势最为显著,此时形状保持特性尤为可取。
- 预报组合无法完全纠正个体预报中严重的系统性误差,凸显了在聚合前优化基础模型的重要性。
- 基于Dropout的集成方法性能劣于通过随机初始化生成的深度集成,表明集成生成方法的选择对最终预报质量有显著影响。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。