Skip to main content
QUICK REVIEW

[论文解读] A theory of maximum likelihood for weighted infection graphs

Justin Khim, Po‐Ling Loh|arXiv (Cornell University)|Jun 13, 2018
Data-Driven Disease Surveillance参考文献 30被引用 4
一句话总结

本文提出了一种最大似然估计框架,用于推断加权感染图中的边权重,其中传播概率取决于累积感染数和边协变量。在对数线性权重模型下,该方法建立了相合性和渐近正态性,利用鞅收敛定理和波利亞瓮过程理论,并为有序与无序传播数据提供了具有理论保证的算法,已在合成数据和埃博拉疫情数据上得到验证。

ABSTRACT

We study the problem of parameter estimation based on infection data from an epidemic outbreak on a graph. We assume that successive infections occur via contagion; i.e., transmissions can only spread across existing directed edges in the graph. Our stochastic spreading model allows individual nodes to be infected more than once, and the probability of the transmission spreading across a particular edge is proportional to both the cumulative number of times the source nodes has been infected in previous stages of the epidemic and the weight parameter of the edge. We propose a maximum likelihood estimator for inferring the unknown edge weights when full information is available concerning the order and identity of successive edge transmissions. When the weights take a particular form as exponential functions of a linear combination of known edge covariates, we show that maximum likelihood estimation amounts to optimizing a convex function, and produces a solution that is both consistent and asymptotically normal. Our proofs are based on martingale convergence theorems and the theory of weighted Pólya urns. We also show how our theory may be generalized to settings where the weights are not exponential. Finally, we analyze the case where the available infection data comes in the form of an unordered set of edge transmissions. We propose two algorithms for weight parameter estimation in this setting and derive corresponding theoretical guarantees. Our methods are validated using both synthetic data and real-world data from the Ebola spread in West Africa.

研究动机与目标

  • 开发一种基于流行病传播数据的统计框架,用于估计加权感染图中未知的边权重。
  • 将传播概率建模为源节点感染次数和边权重的函数,允许重复感染。
  • 在对数线性权重模型下,为最大似然估计提供理论保证——相合性和渐近正态性。
  • 通过两种新提出的算法,将该框架扩展至无序传播数据场景。
  • 在真实世界埃博拉传播数据和合成网络上验证该方法,识别出影响传播的关键协变量。

提出的方法

  • 构建一个随机传播模型,其中边上的传播概率取决于源节点的累积感染数和边的权重参数。
  • 当完整传播序列被观测到时,使用最大似然估计(MLE)推断边权重。
  • 假设边权重为已知边协变量线性组合的指数函数,从而将MLE简化为凸优化问题。
  • 利用鞅收敛定理和加权波利亞瓮过程,证明MLE的相合性和渐近正态性。
  • 针对无序传播情况,提出两种算法:一种基于迭代重加权,另一种基于代理似然函数。
  • 在较弱的正则性条件下,为两种算法推导出理论收敛保证。

实验结果

研究问题

  • RQ1哪些边协变量对网络中感染传播的影响最为显著?
  • RQ2在存在重复感染的加权感染图中,最大似然估计能否一致地恢复边权重?
  • RQ3当仅观测到无序传播集合时,如何进行统计推断?
  • RQ4在非独立同分布的流行病设定下,MLE可建立哪些理论性质,如相合性和渐近正态性?
  • RQ5基于人口密度、旅行时间及边境等协变量的边权重如何影响疾病传播?埃博拉真实数据中已观察到此类影响。

主要发现

  • 当边权重被建模为边协变量线性组合的指数函数时,边权重的MLE具有相合性和渐近正态性。
  • 在对数线性权重模型下,对数似然函数的凸性使得优化和理论分析均得以高效实现。
  • 在埃博拉数据分析中,区域间距离以及源区域和目标区域的人口规模是最具影响力的预测变量,t统计量超过1,000。
  • 国际边境具有强烈的正向影响(系数 = 3.027),而特定国家间跨境传播(如利比里亚至塞拉利昂)则表现出显著负向系数(如利比里亚至塞拉利昂的系数为 -3.866)。
  • 模型识别出共同语言和人口密度为显著协变量,其标准化系数分别为0.845和0.312。
  • 针对无序数据提出的两种算法实现了理论收敛性,在合成数据和真实世界数据上表现良好,展现出对数据结构的强鲁棒性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。