Skip to main content
QUICK REVIEW

[论文解读] On the predictability of infectious disease outbreaks

Samuel V. Scarpino, Giovanni Petri|arXiv (Cornell University)|Mar 21, 2017
Evolution and Genetic Dynamics参考文献 54被引用 4
一句话总结

本研究利用排列熵作为模型无关的可预测性度量,探究了10种历史传染病在时间序列上的基本预测极限。研究发现存在一个普遍存在的熵屏障,超过该屏障后准确预测将变得不可能;然而,对于大多数疾病而言,该屏障远超单次疫情的时间尺度,意味着短期至中期的可靠预测是可行的——尤其当动态、数据驱动的模型能够考虑传播动力学的演变和社交网络异质性时。

ABSTRACT

Infectious disease outbreaks recapitulate biology: they emerge from the multi-level interaction of hosts, pathogens, and their shared environment. As a result, predicting when, where, and how far diseases will spread requires a complex systems approach to modeling. Recent studies have demonstrated that predicting different components of outbreaks--e.g., the expected number of cases, pace and tempo of cases needing treatment, demand for prophylactic equipment, importation probability etc.--is feasible. Therefore, advancing both the science and practice of disease forecasting now requires testing for the presence of fundamental limits to outbreak prediction. To investigate the question of outbreak prediction, we study the information theoretic limits to forecasting across a broad set of infectious diseases using permutation entropy as a model independent measure of predictability. Studying the predictability of a diverse collection of historical outbreaks--including, chlamydia, dengue, gonorrhea, hepatitis A, influenza, measles, mumps, polio, and whooping cough--we identify a fundamental entropy barrier for infectious disease time series forecasting. However, we find that for most diseases this barrier to prediction is often well beyond the time scale of single outbreaks. We also find that the forecast horizon varies by disease and demonstrate that both shifting model structures and social network heterogeneity are the most likely mechanisms for the observed differences across contagions. Our results highlight the importance of moving beyond time series forecasting, by embracing dynamic modeling approaches, and suggest challenges for performing model selection across long time series. We further anticipate that our findings will contribute to the rapidly growing field of epidemiological forecasting and may relate more broadly to the predictability of complex adaptive systems.

研究动机与目标

  • 确定是否存在与模型质量或数据可用性无关的、预测传染病疫情的内在基本限制。
  • 评估疾病时间序列的可预测性是否受信息论熵所衡量的系统内在复杂性的约束。
  • 评估长期预测精度的下降是否并非由于数据稀缺,而是由于传播模式的结构性和动态变化。
  • 研究社交网络异质性和传播动力学变化等因素如何影响不同疾病之间的可预测性。
  • 挑战长期时间序列数据总是能提高预测精度的假设,特别是在复杂、自适应的社会生物系统中。

提出的方法

  • 在10种历史传染病时间序列(如麻疹、流感、登革热)中应用排列熵——一种无需模型的信息论时间序列复杂性和可预测性度量。
  • 使用52周滑动窗口计算时间序列上的排列熵,识别现实数据中可预测性高低的时期。
  • 将观测到的熵值与随机过程(白噪声)和确定性过程(带噪声的正弦波)的参考值进行比较,以建立基线可预测性水平。
  • 结合生物学和流行病学知识,利用模拟疫情探讨传播动力学变化(如疫苗覆盖率、二代感染率)对可预测性的影响。
  • 评估数据长度对可预测性的影响,检验更长的时间序列是否始终能提升预测性能。
  • 采用交叉验证和模型选择技术,评估长时间序列是否能可靠地支持正确模型结构的选择,检验模型选择中是否存在“无免费午餐”现象。

实验结果

研究问题

  • RQ1是否存在一个与建模或数据质量无关的信息论基本极限,限制传染病疫情的可预测性?
  • RQ2疾病时间序列的可预测性如何随时间变化?哪些因素驱动可预测性的波动?
  • RQ3传播动力学的变化(如社交网络结构或疫苗覆盖率的变化)在多大程度上影响预测精度?
  • RQ4增加时间序列数据长度是否总是能提高预测性能?在复杂系统中,是否可能引发虚假可预测性?
  • RQ5动态、数据驱动的模型能否克服时间序列预测中识别出的熵屏障?这对流行病学中的模型选择意味着什么?

主要发现

  • 传染病时间序列中存在一个普遍的熵屏障,代表了可预测性的基本极限,超过该极限后预测在统计上变得不可行。
  • 对于大多数研究的疾病(包括麻疹、流感和登革热),熵屏障远超单次疫情的典型持续时间,表明短期至中期的可靠预测通常是可以实现的。
  • 真实疾病数据中的排列熵值随时间显著波动,存在可预测性高的时期(类似确定性信号)和可预测性低的时期(类似白噪声),表明预测精度具有时间依赖性。
  • 增加数据长度并不总是能提升可预测性;在某些情况下,更长的时间序列反而因传播动力学的结构性变化(如接触网络或免疫力水平的变化)导致可预测性降低。
  • 随着数据跨度的延长,模型性能下降表明最优模型结构可能随时间与尺度而变化,支持采用自适应、迭代校准的模型。
  • 研究结果表明,在传染病预测的模型选择中存在“无免费午餐”现象,即没有一种模型结构能在所有时间窗口中始终优于其他模型,尤其是在非平稳系统中。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。