Skip to main content
QUICK REVIEW

[Paper Review] The graphical structure of respondent-driven sampling

Forrest W. Crawford|arXiv (Cornell University)|Jun 3, 2014
HIV, Drug Use, Sexual Risk27 references4 citations
TL;DR

This paper proposes a continuous-time stochastic model for respondent-driven sampling (RDS) that leverages recruitment timing, coupon use patterns, and respondents' network degrees to infer the hidden social network structure. By modeling recruitment as a time-continuous process and framing the resulting subgraph as an exponential random graph, the method enables computationally efficient estimation of the underlying network, validated through simulations and applied to an RDS study of injection drug users in St. Petersburg, Russia.

ABSTRACT

Respondent-driven sampling (RDS) is a chain-referral method for sampling members of a hidden or hard-to-reach population such as sex workers, homeless people, or drug users via their social network. Most methodological work on RDS has focused on inference of population means under the assumption that subjects' network degree determines their probability of being sampled. Criticism of existing estimators is usually focused on missing data: the underlying network is only partially observed, so it is difficult to determine correct sampling probabilities. In this paper, we show that data collected in ordinary RDS studies contain information about the structure of the respondents' social network. We construct a continuous-time model of RDS recruitment that incorporates the time series of recruitment events, the pattern of coupon use, and the network degrees of sampled subjects. Together, the observed data and the recruitment model place a well-defined probability distribution on the recruitment-induced subgraph of respondents. We show that this distribution can be interpreted as an exponential random graph model and develop a computationally efficient method for estimating the hidden graph. We validate the method using simulated data and apply the technique to an RDS study of injection drug users in St. Petersburg, Russia.

Motivation & Objective

  • To address the limitations of existing RDS estimators that rely on incomplete network data and assume sampling probabilities depend solely on network degree.
  • To develop a method that extracts structural information about the hidden social network from RDS data, which is typically treated as missing or unobserved.
  • To model the recruitment process as a continuous-time stochastic process incorporating time series of recruitment events, coupon use, and network degrees.
  • To estimate the hidden network structure using a computationally efficient framework based on exponential random graph models.
  • To validate the method’s robustness under model mis-specification, particularly in the waiting time distribution of recruitment events.

Proposed method

  • Formulates a continuous-time Markov process for RDS recruitment, where each edge in the hidden network has an independent exponential waiting time for recruitment.
  • Introduces compatibility conditions (Definition 4) that constrain the possible structure of the recruitment-induced subgraph based on observed recruitment times, degrees, and coupon use.
  • Models the recruitment-induced subgraph as an exponential random graph, where edge probabilities depend on the rate parameter λ and the number of susceptible edges.
  • Uses a Bayesian framework with a Gamma prior for λ, informed by bounds derived from observed data to ensure prior plausibility.
  • Applies a Metropolis-Hastings algorithm to sample from the posterior distribution of the hidden graph, enabling inference of network structure.
  • Employs empirical Bayes methods to calibrate the prior on λ using observed recruitment patterns and degree distributions in real data.

Experimental results

Research questions

  • RQ1Can the recruitment process in RDS be modeled as a continuous-time stochastic process that captures temporal dynamics and network structure?
  • RQ2To what extent can the hidden social network structure be recovered from partial observations of recruitment times, degrees, and coupon use?
  • RQ3How robust is the network inference method when the assumed waiting time distribution for recruitment is mis-specified?
  • RQ4Does the proposed model improve upon existing RDS estimators by incorporating network topology and temporal dynamics?
  • RQ5Can the method accurately estimate the edge-wise recruitment rate λ and reconstruct the true network topology in real-world RDS data?

Key findings

  • The method successfully reconstructs the hidden network structure with high accuracy, achieving a mean accuracy of 0.974 under correctly specified exponential waiting times.
  • True positive rate (TPR) and true negative rate (TNR) remain high (0.974 and 0.990, respectively) under correct model specification, indicating strong edge detection performance.
  • Under mis-specification of the waiting time distribution (e.g., Gamma with δ < 1 or δ > 1), the method remains robust in subgraph reconstruction, though λ estimates show bias: overestimated when δ < 1 and underestimated when δ > 1.
  • The estimated edge-wise recruitment rate λ was found to be biased under model mis-specification, with mean estimates deviating from the true λ = 1 by up to 0.10 (for δ = 0.1) and 0.01 (for δ = 0.01).
  • The empirical Bayesian prior for λ, bounded between 9.8×10⁻⁴ and 4.2×10⁻² based on St. Petersburg data, ensures a plausible and informative prior distribution.
  • The method demonstrates strong performance in both simulated data and real-world RDS data from an injection drug user study in St. Petersburg, Russia, with high reconstruction accuracy and stable estimation of λ.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.