Skip to main content
QUICK REVIEW

[Paper Review] The Sensitivity of Respondent-driven Sampling Method

Xin Lü, Linus Bengtsson|arXiv (Cornell University)|Feb 11, 2010
HIV, Drug Use, Sexual Risk21 references16 citations
TL;DR

This paper evaluates the sensitivity of Respondent-Driven Sampling (RDS) by simulating violations of its core assumptions using a large LGBT online community network. Results show that RDS estimates are highly biased when networks are directed or recruitment is non-random based on outcome-related traits, but remain robust under low response rates and reporting errors if these two key issues are absent.

ABSTRACT

Researchers in many scientific fields make inferences from individuals to larger groups. For many groups however, there is no list of members from which to take a random sample. Respondent-driven sampling (RDS) is a relatively new sampling methodology that circumvents this difficulty by using the social networks of the groups under study. The RDS method has been shown to provide unbiased estimates of population proportions given certain conditions. The method is now widely used in the study of HIV-related high-risk populations globally. In this paper, we test the RDS methodology by simulating RDS studies on the social networks of a large LGBT web community. The robustness of the RDS method is tested by violating, one by one, the conditions under which the method provides unbiased estimates. Results reveal that the risk of bias is large if networks are directed, or respondents choose to invite persons based on characteristics that are correlated with the study outcomes. If these two problems are absent, the RDS method shows strong resistance to low response rates and certain errors in the participants' reporting of their network sizes. Other issues that might affect the RDS estimates, such as the method for choosing initial participants, the maximum number of recruitments per participant, sampling with or without replacement and variations in network structures, are also simulated and discussed.

Motivation & Objective

  • To assess the robustness of RDS estimators under violations of their underlying assumptions.
  • To investigate how network structure, homophily, and recruitment behaviors affect RDS estimation accuracy.
  • To evaluate the impact of practical implementation choices such as seed selection, coupon distribution, and sampling with/without replacement.
  • To determine under which conditions RDS provides unbiased estimates in real-world settings.
  • To simulate and quantify biases arising from non-random recruitment, directed edges, and inaccurate degree reporting.

Proposed method

  • Simulated RDS studies on a large, real-world LGBT web community network to test estimator performance.
  • Violated one assumption at a time—such as reciprocity, connectedness, sampling with replacement, degree reporting accuracy, and random recruitment—while measuring bias and error.
  • Used the RDSII estimator: $ \hat{P}_A = \frac{\sum_{i \in A \cap S} d_i^{-1}}{\sum_{i \in S} d_i^{-1}} $, where $ d_i $ is the personal network size of individual $ i $.
  • Compared results across different network types: original, edge-added (higher average degree), and edge-randomized (rewired edges).
  • Varied key design parameters: number of seeds (1–10), number of coupons (1–3), sampling with or without replacement.
  • Modeled non-random recruitment by introducing probabilities dependent on individual characteristics (e.g., homophily) or independent of them.

Experimental results

Research questions

  • RQ1How does network directionality affect the bias and precision of RDS estimates?
  • RQ2What is the impact of non-random recruitment based on characteristics correlated with study outcomes?
  • RQ3How sensitive is RDS to inaccurate reporting of personal network size?
  • RQ4How do sampling with or without replacement affect estimator performance?
  • RQ5How do network structure and homophily influence the accuracy of RDSII estimations?

Key findings

  • RDS estimates exhibit large bias when networks are directed, as the assumption of reciprocal (undirected) relationships is violated.
  • Significant bias arises when recruitment is non-random and based on characteristics correlated with the outcome, such as homophily or social clustering.
  • When both directed networks and non-random recruitment are absent, RDS shows strong resistance to low response rates and errors in self-reported network sizes.
  • The RDSII estimator remains highly accurate under proper conditions: with undirected networks and random recruitment, bias was less than 0.001 even at sample sizes of 500.
  • Sampling without replacement increases bias compared to with replacement, especially in sparse or skewed networks.
  • Estimation performance deteriorates in networks with high skew in degree distribution and high homophily, though performance improves when homophily is low.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.