[论文解读] Time Series Methods and Ensemble Models to Nowcast Dengue at the State Level in Brazil
本研究开发并评估了时间序列与集成模型,利用整合的数据流——临床监测、天气、卫星图像和互联网搜索趋势——对巴西各州登革热发病率进行实时预测。模型表现出高精度,25个州的皮尔逊相关系数超过80%,27个州的中位相关系数达91.75%,表明外生数据显著提升了预测性能,且集成方法增强了实时公共卫生决策中预测的稳健性与可靠性。
Predicting an infectious disease can help reduce its impact by advising public health interventions and personal preventive measures. Novel data streams, such as Internet and social media data, have recently been reported to benefit infectious disease prediction. As a case study of dengue in Brazil, we have combined multiple traditional and non-traditional, heterogeneous data streams (satellite imagery, Internet, weather, and clinical surveillance data) across its 27 states on a weekly basis over seven years. For each state, we nowcast dengue based on several time series models, which vary in complexity and inclusion of exogenous data. The top-performing model varies by state, motivating our consideration of ensemble approaches to automatically combine these models for better outcomes at the state level. Model comparisons suggest that predictions often improve with the addition of exogenous data, although similar performance can be attained by including only one exogenous data stream (either weather data or the novel satellite data) rather than combining all of them. Our results demonstrate that Brazil can be nowcasted at the state level with high accuracy and confidence, inform the utility of each individual data stream, and reveal potential geographic contributors to predictive performance. Our work can be extended to other spatial levels of Brazil, vector-borne diseases, and countries, so that the spread of infectious disease can be more effectively curbed.
研究动机与目标
- 通过整合临床监测、天气、卫星图像和互联网搜索趋势等多样化数据流,提升巴西实时登革热预测的准确性。
- 评估新型数据流(如卫星与互联网数据)相对于传统临床监测和天气数据的预测贡献。
- 开发并验证集成建模方法,通过结合多种时间序列模型以增强各州层面预测的稳健性与准确性。
- 识别与巴西各州模型性能差异相关的地理与社会经济因素。
- 为利用异构数据源与集成建模技术,在巴西及其他类似国家推广登革热等媒介传播疾病实时预测提供可迁移的框架。
提出的方法
- 将巴西卫生部提供的每周登革热病例数与外生数据相结合:基于卫星的植被指数与温度指数、天气数据以及互联网搜索查询量。
- 在每个州应用四种时间序列模型:SARIMA、VAR、基于局部回归(LOESS)的STL分解及其变体,每种模型均整合不同子集的外生变量。
- 采用截尾均值集成方法,通过剔除预测值中上下20%的极端值来减少异常值影响,整合各模型的预测结果。
- 采用加权均值集成方法,模型权重基于验证集上的表现确定——权重与各模型在验证集中最小化绝对误差的频率成正比。
- 保守计算95%预测区间,取截尾模型集合中预测区间上下限的最小值与最大值。
- 在2015–2016年测试窗口上验证模型,主要评估指标为皮尔逊相关系数与预测区间的经验覆盖度。
实验结果
研究问题
- RQ1在巴西各州层面,传统与非传统数据流(如卫星、天气、互联网)的不同组合如何影响登革热实时预测的准确性?
- RQ2在不同巴西州中,哪些时间序列模型在登革热实时预测中表现最佳?最优模型是否存在空间差异?
- RQ3与单一模型相比,集成模型在实时登革热预测中在多大程度上提升了预测性能与可靠性?
- RQ4与巴西各州模型性能差异相关的关键社会经济与地理因素是什么?
- RQ5一个统一的集成框架是否能有效泛化至巴西27个州多样的流行病学与环境条件?
主要发现
- 截尾均值集成模型在2015–2016年测试窗口期间,27个巴西州的拟合值与观测值之间达到91.75%的中位皮尔逊相关系数。
- 25个州(共27个)使用截尾均值集成模型后皮尔逊相关系数超过80%,其中最高达96.44%。
- 一半的州中,95%预测区间的经验覆盖度达到至少96%,表明不确定性量化可靠。
- 引入外生数据——尤其是天气或基于卫星的环境变量——显著提升了模型性能,仅引入一个此类数据流通常即可达到与使用全部数据流相当的性能。
- 模型性能表现出显著的空间自相关性,且与州级教育水平、就业率及人口规模等指标正相关。
- 最优个体模型因州而异,这为采用集成方法实现无需预先知晓各区域最优模型的稳健、州级特定实时预测提供了依据。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。