Skip to main content
QUICK REVIEW

[论文解读] Adaptively stacking ensembles for influenza forecasting with incomplete data

Thomas McAndrew, Nicholas G Reich|arXiv (Cornell University)|Jul 26, 2019
Influenza Virus Research Studies参考文献 53被引用 4
一句话总结

本文提出了一种针对流感样疾病(ILI)的自适应集成预测方法,该方法每周动态更新模型权重,仅使用当季数据,并引入正则化的贝叶斯先验以在噪声大、易被修订的监测数据中稳定预测。该自适应集成方法优于等权重集成模型,且性能与使用多年历史数据训练的静态集成模型相当,为公共卫生响应提供了实用且实时的预测工具。

ABSTRACT

Seasonal influenza infects between 10 and 50 million people in the United States every year, overburdening hospitals during weeks of peak incidence. Named by the CDC as an important tool to fight the damaging effects of these epidemics, accurate forecasts of influenza and influenza-like illness (ILI) forewarn public health officials about when, and where, seasonal influenza outbreaks will hit hardest. Multi-model ensemble forecasts---weighted combinations of component models---have shown positive results in forecasting. Ensemble forecasts of influenza outbreaks have been static, training on all past ILI data at the beginning of a season, generating a set of optimal weights for each model in the ensemble, and keeping the weights constant. We propose an adaptive ensemble forecast that (i) changes model weights week-by-week throughout the influenza season, (ii) only needs the current influenza season's data to make predictions, and (iii) by introducing a prior distribution, shrinks weights toward the reference equal weighting approach and adjusts for observed ILI percentages that are subject to future revisions. We investigate the prior's ability to impact adaptive ensemble performance and, after finding an optimal prior via a cross-validation approach, compare our adaptive ensemble's performance to equal-weighted and static ensembles. Applied to forecasts of short-term ILI incidence at the regional and national level in the US, our adaptive model outperforms a naive equal-weighted ensemble, and has similar or better performance to the static ensemble, which requires multiple years of training data. Adaptive ensembles are able to quickly train and forecast during epidemics, and provide a practical tool to public health officials looking for forecasts that can conform to unique features of a specific season.

研究动机与目标

  • 为解决静态集成预测方法依赖大量历史训练数据和固定权重的局限性,开发一种实时自适应方法,每周更新权重。
  • 在活跃流感季节常见的噪声大、易被修订的ILI监测数据背景下,提升预测的鲁棒性。
  • 评估通过先验分布进行正则化对稳定集成权重和提升预测性能的影响。
  • 在多个季节、地区和预测时域内,比较自适应集成与等权重及静态集成的性能表现。
  • 识别一种可在不同流感季节和预测目标间泛化的最优正则化先验。

提出的方法

  • 该方法使用贝叶斯线性池化方法整合多个分量模型的预测结果,基于近期当季ILI观测数据估计时变权重。
  • 引入先验分布,将集成权重向等权重收缩,以减少对噪声或被修订数据的过拟合。
  • 通过季节间交叉验证优化先验强度,以对数评分作为主要评估指标。
  • 自适应集成模型在每个流感季节内每周进行训练和更新,仅使用当季数据,避免依赖过往季节的数据。
  • 使用随机效应回归模型,比较自适应、等权重和静态集成在不同季节、地区和预测时域内的对数评分差异。
  • 最优先验确定为8%,各季节的第25和第75百分位数分别为6%和8.25%。

实验结果

研究问题

  • RQ1能否通过仅使用当季数据每周更新权重的自适应集成方法,超越使用历史数据训练的静态集成模型?
  • RQ2通过先验分布进行正则化如何影响自适应和静态集成预测的性能?
  • RQ3在面对易被修订的ILI监测数据时,稳定集成权重的最优先验强度是什么?
  • RQ4尽管训练数据显著减少,自适应集成是否能实现与静态集成相当的性能?
  • RQ5是否存在一个一致且可泛化的先验值(例如接近8%),可在不同流感季节和区域中最大化预测准确性?

主要发现

  • 自适应集成在所有季节、地区和预测时域中显著优于等权重集成,2011/2012年季节的对数评分中位数提升达0.13(95%置信区间:0.08, 0.19)。
  • 自适应集成的性能与静态集成相当,多数季节和地区间对数评分的绝对差异很小(例如,β = 0.02,95%置信区间:-0.04, 0.08)。
  • 正则化的最优先验被确定为接近8%,各季节的第25和第75百分位数分别为6%和8.25%。
  • 正则化对自适应和静态集成的性能均有提升,其中自适应模型因数据稀缺和噪声问题,从更大先验中获益更多。
  • 自适应集成在先验为8%时达到峰值性能,且该值在各季节间具有良好的泛化能力,表现为在该值附近对数评分峰值表现一致。
  • 置换检验确认,自适应集成在除2012/2013年外的所有季节中,相对于等权重集成的性能提升均具有统计显著性(p < 0.01)。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。