Skip to main content
QUICK REVIEW

[论文解读] Network Model-Assisted Inference from Respondent-Driven Sampling Data

Krista J. Gile, Mark S. Handcock|arXiv (Cornell University)|Aug 1, 2011
HIV, Drug Use, Sexual Risk参考文献 5被引用 13
一句话总结

本文提出了一种基于模型的响应者驱动抽样(RDS)数据估计方法(μ_MA),通过使用指数族随机图模型(ERGM)作为总体网络的辅助模型,校正初始种子选择和同质性带来的偏差。该方法在初始种子存在偏差时,相比现有RDS估计量显著提升了估计精度,并通过自助法(bootstrap)提供了稳健的不确定性估计,如在乌克兰注射吸毒者群体中HIV患病率估计中的应用所示。

ABSTRACT

Respondent-Driven Sampling is a method to sample hard-to-reach human populations by link-tracing over their social networks. Beginning with a convenience sample, each person sampled is given a small number of uniquely identified coupons to distribute to other members of the target population, making them eligible for enrollment in the study. This can be an effective means to collect large diverse samples from many populations. Inference from such data requires specialized techniques for two reasons. Unlike in standard sampling designs, the sampling process is both partially beyond the control of the researcher, and partially implicitly defined. Therefore, it is not generally possible to directly compute the sampling weights necessary for traditional design-based inference. Any likelihood-based inference requires the modeling of the complex sampling process often beginning with a convenience sample. We introduce a model-assisted approach, resulting in a design-based estimator leveraging a working model for the structure of the population over which sampling is conducted. We demonstrate that the new estimator has improved performance compared to existing estimators and is able to adjust for the bias induced by the selection of the initial sample. We present sensitivity analyses for unknown population sizes and the misspecification of the working network model. We develop a bootstrap procedure to compute measures of uncertainty. We apply the method to the estimation of HIV prevalence in a population of injecting drug users (IDU) in the Ukraine, and show how it can be extended to include application-specific information.

研究动机与目标

  • 本文旨在解决当初始样本为非概率抽样且存在偏差时,RDS数据缺乏可靠推断方法的问题。
  • 旨在校正RDS研究中由于初始种子的便利抽样所引入的选择偏差。
  • 通过引入工作ERGM模型来刻画网络结构,以提升估计精度。
  • 开发一种自助法程序,以量化在总体规模未知及模型误设情况下的不确定性。
  • 展示该方法在整合特定应用的网络特征(如按感染状态的招募模式)方面的灵活性。

提出的方法

  • 所提出的估计量μ_MA是一种基于设计的方法,利用工作ERGM对目标总体的潜在社会网络结构进行建模。
  • 该方法通过引入网络模型,校正由初始种子偏差和同质性引起的不可忽略的选择机制。
  • 通过根据ERGM推导的受访者网络位置进行加权,来校正不等的纳入概率,从而估计总体比例。
  • 开发了自助法程序以计算标准误,同时考虑抽样变异性和模型不确定性。
  • 该方法允许在ERGM框架内纳入辅助信息(如转介模式或特定属性上的同质性)。
  • 该方法应用于乌克兰注射吸毒者群体的真实RDS数据,并对总体规模和网络模型误设进行了敏感性分析。

实验结果

研究问题

  • RQ1基于模型的估计量在多大程度上能减少RDS研究中由初始种子选择引起的偏差?
  • RQ2在种子偏差和同质性存在的情况下,所提出的μ_MA估计量与现有RDS估计量(如μ_VH、μ_SS)相比表现如何?
  • RQ3当工作网络模型存在误设或总体规模估计存在误差时,μ_MA估计量的敏感性如何?
  • RQ4该方法能否有效整合特定应用的网络特征,如按感染状态的差异性招募?
  • RQ5网络传递性对估计量性能有何影响?工作模型在高聚类条件下是否仍保持稳健?

主要发现

  • 当初始种子样本存在偏差时,尤其在同质性存在的情况下,μ_MA估计量显著优于现有估计量(如μ_VH和μ_SS)。
  • 该估计量对工作网络模型的适度误设具有鲁棒性,即使在高传递性条件下也仅表现出微小影响。
  • 敏感性分析表明,即使总体规模被中度低估,μ_MA仍表现良好,但严重低估时性能会下降。
  • 自助法程序成功捕捉了估计的不确定性,为推断提供了可靠的标误差。
  • 在乌克兰注射吸毒者群体的HIV患病率应用中,该方法的估计结果与更稳健的μ_SH估计量接近,表明其偏差校正能力得到改善。
  • 该方法能够灵活整合辅助数据(如转介模式),增强模型现实性与估计精度。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。