Skip to main content
QUICK REVIEW

[Paper Review] Central limit theorems for network driven sampling

Li Xiao, Karl Rohe|arXiv (Cornell University)|Sep 15, 2015
HIV, Drug Use, Sexual Risk20 references3 citations
TL;DR

This paper establishes central limit theorems for network-driven sampling, modeling Respondent-Driven Sampling as a tree-indexed Markov process. It proves the Volz-Heckathorn estimator is asymptotically normal under a critical threshold and shows that in certain cases, the sample average can have lower mean squared error than inverse probability weighting due to reduced variance when outcomes are uncorrelated with sampling weights.

ABSTRACT

Respondent-Driven Sampling is a popular technique for sampling hidden populations. This paper models Respondent-Driven Sampling as a Markov process indexed by a tree. Our main results show that the Volz-Heckathorn estimator is asymptotically normal below a critical threshold. The key technical difficulties stem from (i) the dependence between samples and (ii) the tree structure which characterizes the dependence. The theorems allow the growth rate of the tree to exceed one and suggest that this growth rate should not be too large. To illustrate the usefulness of these results beyond their obvious use, an example shows that in certain cases the sample average is preferable to inverse probability weighting. We provide a test statistic to distinguish between these two cases.

Motivation & Objective

  • To establish asymptotic normality for estimators in Respondent-Driven Sampling (RDS) under a tree-indexed Markov process model.
  • To address the challenges of dependent samples and complex tree-structured referral networks in RDS.
  • To determine when the sample average outperforms inverse probability weighting (IPW) in terms of mean squared error.
  • To provide a test statistic for selecting between the sample average and IPW estimator based on bias and variance trade-offs.

Proposed method

  • Models RDS as a tree-indexed Markov process on a finite, undirected graph, with referrals forming a rooted tree structure.
  • Uses the Volz-Heckathorn estimator as a proxy for the inverse probability weighted (IPW) estimator, enabling asymptotic analysis.
  • Applies central limit theorems to the sample average and Volz-Heckathorn estimator under general tree growth rates.
  • Derives variance expressions for both estimators using graph-theoretic quantities, including transition probabilities and stationary measures.
  • Introduces a test statistic to assess whether the bias of the sample average is zero, enabling selection between sample average and IPW.
  • Employs spectral graph theory and eigenvalue decomposition of the transition matrix to analyze dependence structure and variance components.

Experimental results

Research questions

  • RQ1Under what conditions is the Volz-Heckathorn estimator asymptotically normal in network-driven sampling?
  • RQ2When does the sample average have lower mean squared error than the inverse probability weighted estimator in RDS?
  • RQ3How does the growth rate of the referral tree affect the asymptotic variance of estimators in RDS?
  • RQ4What is the role of the correlation between outcomes and sampling weights in determining estimator efficiency?
  • RQ5Can a test statistic be constructed to distinguish between cases where the sample average is preferable to IPW?

Key findings

  • The Volz-Heckathorn estimator is asymptotically normal when the tree growth rate is below a critical threshold.
  • The sample average can have lower mean squared error than the IPW estimator when outcomes are uncorrelated with sampling weights.
  • The variance of the IPW estimator exceeds that of the sample average by at least a term proportional to $ \frac{C_2}{N C_1^2} $, indicating potential inefficiency of IPW.
  • A test statistic is derived to assess the null hypothesis that the sample average is unbiased, enabling data-driven selection between estimators.
  • Simulations show that $ \mathbb{G} $, the function governing variance, is often convex but not always, even under regular offspring distributions.
  • The theoretical framework provides a good approximation for real-world RDS even when the i.i.d. assumption is violated, as validated by simulations.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.