Skip to main content
QUICK REVIEW

[论文解读] Identification of homophily and preferential recruitment in respondent-driven sampling

Forrest W. Crawford, Peter M. Aronow|arXiv (Cornell University)|Nov 17, 2015
HIV, Drug Use, Sexual Risk参考文献 1被引用 4
一句话总结

本文严格定义了受访者驱动抽样(RDS)中的同质性(homophily)与偏好招募(preferential recruitment),并证明仅凭RDS数据无法对二者进行点识别,原因在于未观测到的网络结构。通过非参数识别区域与随机优化方法,作者表明:在缺乏额外网络数据的情况下,RDS研究中关于同质性或招募偏差的实证主张在统计上不可靠。

ABSTRACT

Respondent-driven sampling (RDS) is a link-tracing procedure for surveying hidden or hard-to-reach populations in which subjects recruit other subjects via their social network. There is significant research interest in detecting clustering or dependence of epidemiological traits in networks, but researchers disagree about whether data from RDS studies can reveal it. Two distinct mechanisms account for dependence in traits of recruiters and recruitees in an RDS study: homophily, the tendency for individuals to share social ties with others exhibiting similar characteristics, and preferential recruitment, in which recruiters do not recruit uniformly at random from their available alters. The different effects of network homophily and preferential recruitment in RDS studies have been a source of confusion in methodological research on RDS, and in empirical studies of the social context of health risk in hidden populations. In this paper, we give rigorous definitions of homophily and preferential recruitment and show that neither can be measured precisely in general RDS studies. We derive nonparametric identification regions for homophily and preferential recruitment and show that these parameters are not point identified unless the network takes a degenerate form. The results indicate that claims of homophily or recruitment bias measured from empirical RDS studies may not be credible. We apply our identification results to a study involving both a network census and RDS on a population of injection drug users in Hartford, CT.

研究动机与目标

  • 澄清RDS中同质性与偏好招募的独立机制,二者在方法论与实证研究中常被混淆。
  • 证明由于未观测到的网络边的存在,仅凭标准RDS数据无法对同质性或偏好招募实现点识别。
  • 在一般RDS条件下,推导出这两个参数的非参数识别区域。
  • 利用美国康涅狄格州哈特福德市的真实数据集,评估RDS研究中关于同质性或招募偏差的实证主张的可信度。
  • 为评估RDS在检测流行病学特征网络聚类方面局限性,提供理论基础。

提出的方法

  • 作者将同质性定义为具有相似特征的个体倾向于形成社会联系,将偏好招募定义为从网络邻居中非均匀选择招募对象。
  • 将问题形式化为非参数识别挑战,表明存在多种未观测到的网络配置可产生相同的RDS数据。
  • 使用随机优化与马尔可夫链蒙特卡洛(MCMC)抽样方法,探索与观测到的RDS数据兼容的增强子图与特征集合空间,以估计识别区域。
  • 提出一种两阶段MCMC提议机制:(1) 通过引入新顶点进行边的增删,(2) 对未采样顶点的特征值进行切换。
  • 该方法依赖Hajek(1988)提出的冷却计划与收敛准则,以确保MCMC链渐近地收敛至适应度函数的全局最大值。
  • 通过探索与观测到的RDS数据及招募度数兼容的所有可能网络与特征配置集合,推导出识别区域。

实验结果

研究问题

  • RQ1仅凭受访者驱动抽样数据,能否对潜在社会网络中的同质性实现点识别?
  • RQ2在未观测到完整网络结构的情况下,能否精确测量RDS中的偏好招募行为?
  • RQ3RDS研究中关于同质性或招募偏差的实证主张,在多大程度上依赖于不可验证的假设?
  • RQ4在一般RDS条件下,同质性与偏好招募的非参数识别区域是什么?
  • RQ5理论上的识别限制如何影响RDS数据在公共卫生与流行病学建模中的解释?

主要发现

  • 由于存在多种未观测到的网络配置可产生完全相同的观测招募数据,因此在一般RDS研究中,同质性与偏好招募无法实现点识别。
  • 这两个参数的识别区域是非退化的,这意味着即使拥有完整的RDS数据,若无额外网络信息,也无法精确测量同质性或招募偏差。
  • 作者证明,在特定冷却计划下,MCMC链收敛至适应度函数全局最大值的概率趋近于1,从而确保渐近收敛至正确的识别区域。
  • 在对美国康涅狄格州哈特福德市注射吸毒者群体的真实世界应用中,同质性与偏好招募的识别区域范围较宽,表明实证估计的精度有限。
  • 本研究结论认为,仅基于RDS数据而声称存在同质性或招募偏差,在缺乏外部网络数据的情况下,其统计可信度不足。
  • 研究结果挑战了利用RDS数据估计疾病传播模型中 assortative mixing(同配性混合)或网络聚类有效性的合理性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。