Skip to main content
QUICK REVIEW

[Paper Review] Estimating Hidden Population Size using Respondent-Driven Sampling Data

Mark S. Handcock, Krista J. Gile|arXiv (Cornell University)|Sep 27, 2012
HIV, Drug Use, Sexual Risk5 references4 citations
TL;DR

This paper proposes a Bayesian method to estimate the size of hidden human populations using only Respondent-Driven Sampling (RDS) data, leveraging the decreasing sequence of personal network sizes observed during sampling. The approach uses a successive sampling approximation to model population depletion, yielding reliable population size estimates with good frequentist coverage and improved inference for aggregate characteristics.

ABSTRACT

Respondent-Driven Sampling (RDS) is an approach to sampling design and inference in hard-to-reach human populations. Typically, a sampling frame is not available, and population members are difficult to identify or recruit from broader sampling frames. Common examples include injecting drug users, men who have sex with men, and female sex workers. Most analysis of RDS data has focused on estimating aggregate characteristics, such as disease prevalence. However, RDS is often conducted in settings where the population size is unknown and of great independent interest. This paper presents an approach to estimating the size of a target population based on data collected through RDS. The proposed approach uses a successive sampling approximation to RDS to leverage information in the ordered sequence of observed personal network sizes. The inference uses the Bayesian framework, allowing for the incorporation of prior knowledge. A flexible class of priors for the population size is proposed that aids elicitation. An extensive simulation study provides insight into the performance of the method for estimating population size under a broad range of conditions. A further study shows the approach also improves estimation of aggregate characteristics. A particular choice of the prior produces interval estimates with good frequentist properties. Finally, the method demonstrates sensible results when used to estimate the numbers of sub-populations most at risk for HIV in two cities in El Salvador.

Motivation & Objective

  • To develop a method for estimating the size of hard-to-reach human populations using only RDS data, without requiring additional data sources.
  • To exploit the ordered sequence of personal network sizes in RDS to infer population size, treating sampling dependence as informative rather than a nuisance.
  • To provide a flexible Bayesian framework that incorporates prior knowledge and enables coherent combination with other estimation methods.
  • To improve estimation of aggregate characteristics (e.g., disease prevalence) by using more accurate population size estimates.
  • To demonstrate the method’s performance across diverse sampling conditions and real-world settings, including HIV risk groups in El Salvador.

Proposed method

  • Uses a successive sampling approximation to model the RDS process, where larger network sizes are more likely to be sampled earlier, reflecting population depletion.
  • Applies a Bayesian hierarchical model to estimate population size N, with a flexible class of priors to aid elicitation and improve robustness.
  • Models the observed sequence of personal network sizes as a function of sampling order, assuming decreasing sizes indicate population depletion.
  • Employs Markov Chain Monte Carlo (MCMC) for posterior computation, enabling full uncertainty quantification of N.
  • Combines the posterior from this method with other data sources (e.g., capture-recapture, multiplier methods) by using it as a prior in integrated inference.
  • Validates the method through extensive simulation studies across varying population sizes, network structures, and sampling fractions.

Experimental results

Research questions

  • RQ1Can population size be reliably estimated from RDS data alone, without additional data sources?
  • RQ2How does the sequential pattern of observed personal network sizes inform population size estimation under RDS?
  • RQ3What is the performance of the proposed Bayesian method in terms of frequentist coverage and bias across diverse sampling conditions?
  • RQ4Can the estimated population size improve the precision of RDS-based prevalence estimators?
  • RQ5How well does the method perform in real-world applications, such as estimating HIV risk group sizes in El Salvador?

Key findings

  • The method produces interval estimates for population size with good frequentist coverage, even in challenging conditions with strong differential activity.
  • Simulation results show that the method performs well, especially in larger sample fractions and under strong heterogeneity in network sizes.
  • Using the estimated population size in the SS estimator (Gile, 2011) significantly improves the precision of prevalence estimates.
  • The method yields estimates for HIV risk groups in El Salvador that are compatible with UNAIDS guidelines and capture-recapture estimates.
  • The posterior distribution from this method can be used as a prior in multi-method inference, enabling coherent and incremental improvement of population size estimates.
  • In cases with limited information, the method produces wide credible intervals, reflecting genuine uncertainty, which is a more honest assessment than methods that underestimate variability.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.