[论文解读] Data-Centric Epidemic Forecasting: A Survey
本综述提出了一套全面的数据驱动型流行病预测框架,整合了临床、数字、行为、基因组、环境和政策等多种数据源,结合先进的统计学、机器学习及混合建模技术。它识别了在不确定性量化、可解释性以及现实世界部署方面面临的关键挑战,并为支持公共卫生决策提供了可操作、可靠的预测路线图。
The COVID-19 pandemic has brought forth the importance of epidemic forecasting for decision makers in multiple domains, ranging from public health to the economy as a whole. While forecasting epidemic progression is frequently conceptualized as being analogous to weather forecasting, however it has some key differences and remains a non-trivial task. The spread of diseases is subject to multiple confounding factors spanning human behavior, pathogen dynamics, weather and environmental conditions. Research interest has been fueled by the increased availability of rich data sources capturing previously unobservable facets and also due to initiatives from government public health and funding agencies. This has resulted, in particular, in a spate of work on 'data-centered' solutions which have shown potential in enhancing our forecasting capabilities by leveraging non-traditional data sources as well as recent innovations in AI and machine learning. This survey delves into various data-driven methodological and practical advancements and introduces a conceptual framework to navigate through them. First, we enumerate the large number of epidemiological datasets and novel data streams that are relevant to epidemic forecasting, capturing various factors like symptomatic online surveys, retail and commerce, mobility, genomics data and more. Next, we discuss methods and modeling paradigms focusing on the recent data-driven statistical and deep-learning based methods as well as on the novel class of hybrid models that combine domain knowledge of mechanistic models with the effectiveness and flexibility of statistical approaches. We also discuss experiences and challenges that arise in real-world deployment of these forecasting systems including decision-making informed by forecasts. Finally, we highlight some challenges and open problems found across the forecasting pipeline.
研究动机与目标
- 整合疫情暴发以来关于数据驱动型流行病预测的快速增长研究,特别是针对新冠疫情的响应。
- 识别并分类多样化的数据源——包括临床、数字、行为、基因组、环境和政策数据——以提升预测准确性。
- 评估统计模型、机器学习模型以及混合建模方法在流行病预测中的优势与局限性。
- 解决预测系统在不确定性量化、可解释性以及现实世界部署中的关键挑战。
- 提出一个概念性框架,以使预测输出与可操作的公共卫生决策及决策流程相一致。
提出的方法
- 系统性地将20余种数据类型归类至六个领域:临床监测、数字监测、行为数据、基因组、环境和政策数据。
- 将预测任务划分为实值预测、事件预测和流行病学指标估计三类,并采用相应的评估指标。
- 综述统计模型与深度学习模型,包括稀疏线性模型、自回归方法、视觉与语言模型,以及具有时间与空间归纳偏差的神经网络。
- 提出混合建模范式,通过数据同化、参数估计和差异建模,将机制模型(如分 compartment 模型、基于代理的模型)与数据驱动组件相融合。
- 建议采用非参数神经网络模型与不确定性量化技术,以提升校准性,并区分认知不确定性与随机不确定性。
- 倡导采用人机协同系统与技术债务管理,以提升现实世界预测部署中的鲁棒性与可复现性。
实验结果
研究问题
- RQ1非传统数据源(如移动性、社交媒体和废水数据)如何提升流行病预测的准确性?
- RQ2在预测疾病传播时,纯统计模型、深度学习模型与混合机制-统计模型之间的关键方法论差异与权衡是什么?
- RQ3如何更好地量化与传达预测中的不确定性,特别是在基线发病率各异的地区?
- RQ4技术债务与专家干预在现实世界预测系统中扮演何种角色,又该如何系统性地管理?
- RQ5如何通过改进目标设定、可视化手段与评估指标,使预测输出对公共卫生决策者更具可操作性?
主要发现
- 多模态数据的整合——尤其是移动性、在线症状调查与废水监测——在近期的疫情预测挑战中显著提升了短期预测性能。
- 将机制性疾病传播动力学与数据驱动组件相结合的混合模型(例如通过数据同化或差异建模)在许多现实世界预测场景中优于纯统计或纯机制模型。
- 非参数神经网络等模型生成的概率预测相比传统集成方法展现出更优的校准性,尤其在疫情高峰期间对不确定性的捕捉更为准确。
- 不确定性量化仍是重大挑战,覆盖率指标常无法真实反映预测的可靠性,尤其在发病率较低或波动较大的地区。
- 人机协同干预(如数据清洗、参数调优与异常值剔除)构成了一种显著的技术债务形式,影响模型性能,必须在评估中系统性地加以考虑。
- 标准化评估指标(如WIS、MAE与覆盖率)虽被广泛使用,但在不同地区或事件预测(如峰值时间)中可能不够稳健,凸显了开发新型上下文感知评估框架的必要性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。