[Paper Review] Estimating Spread of Contact-Based Contagions in a Population Through Sub-Sampling
This paper proposes PollSpreader and PollSusceptible, novel methods that estimate the spread of contact-based contagions in a population using only sub-sampled individual location data. By directly modeling co-locations between sampled and unobserved individuals through a contact network, the approach provides theoretically grounded upper and lower bounds on disease spread, outperforming scaling and synthetic trajectory methods that fail due to co-location inaccuracies and error amplification over time.
Physical contacts result in the spread of various phenomena such as viruses, gossips, ideas, packages and marketing pamphlets across a population. The spread depends on how people move and co-locate with each other, or their mobility patterns. How far such phenomena spread has significance for both policy making and personal decision making, e.g., studying the spread of COVID-19 under different intervention strategies such as wearing a mask. In practice, mobility patterns of an entire population is never available, and we usually have access to location data of a subset of individuals. In this paper, we formalize and study the problem of estimating the spread of a phenomena in a population, given that we only have access to sub-samples of location visits of some individuals in the population. We show that simple solutions such as estimating the spread in the sub-sample and scaling it to the population, or more sophisticated solutions that rely on modeling location visits of individuals do not perform well in practice, the former because it ignores contacts between unobserved individuals and sampled ones and the latter because it yields inaccurate modeling of co-locations. Instead, we directly model the co-locations between the individuals. We introduce PollSpreader and PollSusceptible, two novel approaches that model the co-locations between individuals using a contact network, and infer the properties of the contact network using the subsample to estimate the spread of the phenomena in the entire population. We show that our estimates provide an upper bound and a lower bound on the spread of the disease in expectation. Finally, using a large high-resolution real-world mobility dataset, we experimentally show that our estimates are accurate, while other methods that do not correctly account for co-locations between individuals result in wrong observations (e.g, premature herd-immunity).
Motivation & Objective
- To address the challenge of estimating contagion spread in a full population when only sub-sampled individual mobility data is available.
- To overcome the limitations of naive scaling and synthetic trajectory generation, which fail to accurately model co-locations between sampled and unobserved individuals.
- To develop a method that directly models co-locations via a contact network to infer population-level spread with theoretical bounds.
- To provide a computationally efficient and accurate alternative to full-population simulations, especially for policy-relevant what-if scenarios.
- To validate the method on real-world mobility data, demonstrating robustness against error amplification seen in existing approaches.
Proposed method
- The method constructs a contact network based on observed co-locations between sampled individuals and infers unobserved co-locations using statistical modeling.
- PollSpreader estimates the upper bound of spread by assuming all unobserved individuals are susceptible to contact with any sampled individual.
- PollSusceptible estimates the lower bound by assuming no additional co-locations beyond those observed in the sample.
- The approach models co-locations directly rather than relying on synthetic trajectory generation or spatial discretization.
- Theoretical analysis proves that PollSpreader provides an upper bound and PollSusceptible a lower bound on the expected spread in the full population.
- The method uses real-world high-resolution mobility data to calibrate and validate the bounds, avoiding reliance on inaccurate synthetic data.
Experimental results
Research questions
- RQ1Can we accurately estimate the spread of a contact-based contagion in a full population using only a sub-sample of mobility data?
- RQ2Why do traditional methods like scaling the sub-sample or generating synthetic trajectories fail in practice?
- RQ3Can we model co-locations between sampled and unobserved individuals directly to improve estimation accuracy?
- RQ4Do the proposed methods provide theoretically justified bounds on the true population-level spread?
- RQ5How does error from incorrect co-location modeling amplify over time in existing methods, and can our approach prevent this?
Key findings
- PollSpreader and PollSusceptible provide theoretically sound upper and lower bounds on the expected spread of a contagion in the full population.
- On the Verily mobility dataset, PollSus_L and PollSus_U closely track the ground-truth spread, while baseline methods like Scale and PollSpreader show premature herd immunity patterns.
- The PollSpreader method incorrectly predicts a decline in infections due to error amplification, even when the true number of infections is increasing.
- Scaling the sub-sample underestimates spread significantly because it ignores co-locations between sampled and unobserved individuals.
- Synthetic trajectory generation methods fail due to the need for sub-meter accuracy over long durations and the resulting data explosion, making large-scale simulations impractical.
- The proposed method outperforms existing approaches by modeling co-locations directly, avoiding the inaccuracies introduced by trajectory synthesis and spatial discretization.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.