[论文解读] Prediction Regions for Poisson and Over-Dispersed Poisson Regression Models with Applications to Forecasting Number of Deaths during the COVID-19 Pandemic
本文提出针对泊松分布和过度离散泊松回归模型的预测区间,用于预测新冠疫情期间的日度和累计死亡人数。该研究引入了一种基于脆弱性(frailty-based)的过度离散泊松模型,以考虑死亡数据中的过度变异,表明在重新分配调整后预测区间更宽,并显示点预测值和不确定性随时间显著增加,尤其是在2020年7月2日之后。
Motivated by the current Coronavirus Disease (COVID-19) pandemic, which is due to the SARS-CoV-2 virus, and the important problem of forecasting daily deaths and cumulative deaths, this paper examines the construction of prediction regions or intervals under the Poisson regression model and for an over-dispersed Poisson regression model. For the Poisson regression model, several prediction regions are developed and their performance are compared through simulation studies. The methods are applied to the problem of forecasting daily and cumulative deaths in the United States (US) due to COVID-19. To examine their performance relative to what actually happened, daily deaths data until May 15th were used to forecast cumulative deaths by June 1st. It was observed that there is over-dispersion in the observed data relative to the Poisson regression model. An over-dispersed Poisson regression model is therefore proposed. This new model builds on frailty ideas in Survival Analysis and over-dispersion is quantified through an additional parameter. The Poisson regression model is a hidden model in this over-dispersed Poisson regression model and obtains as a limiting case when the over-dispersion parameter increases to infinity. A prediction region for the cumulative number of US deaths due to COVID-19 by July 16th, given the data until July 2nd, is presented. Finally, the paper discusses limitations of proposed procedures and mentions open research problems, as well as the dangers and pitfalls when forecasting on a long horizon, with focus on this pandemic where events, both foreseen and unforeseen, could have huge impacts on point predictions and prediction regions.
研究动机与目标
- 为应对新冠疫情期间日度和累计死亡人数预测区间可靠性的关键需求。
- 对报告死亡人数的过度离散化进行建模,这违反了标准泊松分布的假设。
- 通过基于脆弱性(frailty-based)的随机效应开发一种灵活的过度离散泊松回归模型,利用额外参数量化过度离散化。
- 基于截至2020年7月2日的数据,构建截至2020年7月16日的累计死亡人数预测区间。
- 评估数据重新分配和模型敏感性对预测不确定性及点估计的影响。
提出的方法
- 使用DayNum的五阶多项式作为时间趋势,并引入星期几的指示变量以捕捉死亡报告中的周周期模式。
- 提出一种过度离散泊松回归模型,其中过度离散化通过伽马分布的脆弱性项进行建模,当离散参数增大时,泊松模型成为其极限情况。
- 对泊松分布应用正态近似,以在两种模型下构建预测区间。
- 通过模拟研究比较在泊松和过度离散泊松假设下不同预测区间方法的性能。
- 通过组合每日死亡预测并进行累计和传播,构建累计死亡人数的预测区间,同时考虑每日预测中的不确定性。
- 执行数据重新分配调整,以模拟延迟或更正报告对预测准确性和区间宽度的影响。
实验结果
研究问题
- RQ1在泊松和过度离散泊松假设下,不同预测区间方法在日度死亡人数预测中的表现如何?
- RQ2美国报告的新冠疫情死亡数据中过度离散化在多大程度上使标准泊松回归模型失效?
- RQ3数据重新分配调整如何影响累计死亡人数预测区间的宽度和位置?
- RQ4预测不确定性如何随时间演变,尤其是在训练期之后?
- RQ5与标准泊松模型相比,基于脆弱性(frailty-based)的过度离散泊松模型是否能提高疫情预测的可靠性?
主要发现
- 观察到的美国新冠疫情死亡数据相对于泊松模型表现出显著的过度离散化,因此必须采用过度离散模型。
- 在应用数据重新分配调整后,2020年7月16日累计死亡人数的预测区间更宽,表明报告变更导致不确定性增加。
- 在执行重新分配后,2020年7月16日累计死亡人数的点预测值略有提高,反映出数据输入的更新。
- 随着预测时间范围超出2020年7月2日,预测区间显著变宽,反映出长期预测中不确定性的增加。
- 日度死亡人数的预测曲线在2020年7月2日之后呈现上升趋势,与此前的下降趋势形成对比,提示政策变化或行为转变可能产生影响。
- 模型对异常值(特别是纽约州和新泽西州的大规模重新分配事件)的敏感性,凸显了长期预测的局限性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。