[论文解读] Estimating Spread of Contact-Based Contagions in a Population Through Sub-Sampling
本文提出了 PollSpreader 和 PollSusceptible 两种新方法,仅使用子采样个体位置数据即可估计基于接触传播的疾病在人群中的传播情况。通过在接触网络中直接建模采样个体与未观测个体之间的共现行为,该方法提供了理论上可靠的传播上限和下限,优于因共现建模不准确及随时间累积误差而失效的缩放法和合成轨迹法。
Physical contacts result in the spread of various phenomena such as viruses, gossips, ideas, packages and marketing pamphlets across a population. The spread depends on how people move and co-locate with each other, or their mobility patterns. How far such phenomena spread has significance for both policy making and personal decision making, e.g., studying the spread of COVID-19 under different intervention strategies such as wearing a mask. In practice, mobility patterns of an entire population is never available, and we usually have access to location data of a subset of individuals. In this paper, we formalize and study the problem of estimating the spread of a phenomena in a population, given that we only have access to sub-samples of location visits of some individuals in the population. We show that simple solutions such as estimating the spread in the sub-sample and scaling it to the population, or more sophisticated solutions that rely on modeling location visits of individuals do not perform well in practice, the former because it ignores contacts between unobserved individuals and sampled ones and the latter because it yields inaccurate modeling of co-locations. Instead, we directly model the co-locations between the individuals. We introduce PollSpreader and PollSusceptible, two novel approaches that model the co-locations between individuals using a contact network, and infer the properties of the contact network using the subsample to estimate the spread of the phenomena in the entire population. We show that our estimates provide an upper bound and a lower bound on the spread of the disease in expectation. Finally, using a large high-resolution real-world mobility dataset, we experimentally show that our estimates are accurate, while other methods that do not correctly account for co-locations between individuals result in wrong observations (e.g, premature herd-immunity).
研究动机与目标
- 解决仅能获得子采样个体移动数据时,对全人群传播情况估计的挑战。
- 克服朴素缩放法和合成轨迹生成法的局限性,这些方法无法准确建模采样个体与未观测个体之间的共现行为。
- 开发一种通过接触网络直接建模共现行为的方法,以推断全人群传播水平并提供理论边界。
- 为全人群模拟提供一种计算高效且准确的替代方案,尤其适用于政策相关的“如果-那么”情景分析。
- 在真实世界移动数据上验证该方法,证明其对现有方法中常见的误差累积现象具有鲁棒性。
提出的方法
- 该方法基于采样个体之间的观测共现行为构建接触网络,并通过统计建模推断未观测个体的共现行为。
- PollSpreader 通过假设所有未观测个体均可能与任一采样个体接触,来估计传播的上限。
- PollSusceptible 通过假设除样本中观测到的共现外无额外共现,来估计传播的下限。
- 该方法直接建模共现行为,而非依赖于合成轨迹生成或空间离散化。
- 理论分析证明,PollSpreader 提供全人群预期传播的上界,PollSusceptible 提供下界。
- 该方法使用真实世界高分辨率移动数据校准并验证边界,避免依赖不准确的合成数据。
实验结果
研究问题
- RQ1我们能否仅使用子采样移动数据,准确估计全人群中基于接触传播的疾病传播情况?
- RQ2为何传统方法如缩放子样本或生成合成轨迹在实践中会失效?
- RQ3我们能否直接建模采样个体与未观测个体之间的共现行为,以提高估计精度?
- RQ4所提出的方法能否提供对真实全人群传播水平的理论支持边界?
- RQ5现有方法中,错误共现建模导致的误差如何随时间累积?我们的方法能否防止此类累积?
主要发现
- PollSpreader 和 PollSusceptible 为全人群中疾病传播的预期值提供了理论可靠的上下界。
- 在 Verily 移动数据集上,PollSus_L 和 PollSus_U 紧密跟踪真实传播情况,而 Scale 和 PollSpreader 等基线方法则表现出过早群体免疫的模式。
- PollSpreader 方法因误差累积错误预测了感染数下降,即使真实感染数仍在上升。
- 缩放子样本会显著低估传播,因为它忽略了采样个体与未观测个体之间的共现行为。
- 合成轨迹生成方法因需要长时间保持亚米级精度,且导致数据量爆炸,使大规模模拟变得不切实际。
- 所提方法通过直接建模共现行为,避免了轨迹合成与空间离散化带来的不准确性,因此优于现有方法。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。