Skip to main content
QUICK REVIEW

[论文解读] The graphical structure of respondent-driven sampling

Forrest W. Crawford|arXiv (Cornell University)|Jun 3, 2014
HIV, Drug Use, Sexual Risk参考文献 27被引用 4
一句话总结

本文提出了一种连续时间随机模型,用于响应者驱动抽样(RDS)中的隐藏社交网络结构推断,该模型利用招募时间、代金券使用模式以及受访者网络度数。通过将招募过程建模为连续时间过程,并将所得子图视为指数随机图,该方法实现了对潜在网络的计算高效估计,其有效性通过模拟实验得到验证,并应用于俄罗斯圣彼得堡注射吸毒者群体的RDS研究。

ABSTRACT

Respondent-driven sampling (RDS) is a chain-referral method for sampling members of a hidden or hard-to-reach population such as sex workers, homeless people, or drug users via their social network. Most methodological work on RDS has focused on inference of population means under the assumption that subjects' network degree determines their probability of being sampled. Criticism of existing estimators is usually focused on missing data: the underlying network is only partially observed, so it is difficult to determine correct sampling probabilities. In this paper, we show that data collected in ordinary RDS studies contain information about the structure of the respondents' social network. We construct a continuous-time model of RDS recruitment that incorporates the time series of recruitment events, the pattern of coupon use, and the network degrees of sampled subjects. Together, the observed data and the recruitment model place a well-defined probability distribution on the recruitment-induced subgraph of respondents. We show that this distribution can be interpreted as an exponential random graph model and develop a computationally efficient method for estimating the hidden graph. We validate the method using simulated data and apply the technique to an RDS study of injection drug users in St. Petersburg, Russia.

研究动机与目标

  • 解决现有RDS估计器的局限性,这些估计器依赖于不完整的网络数据,并假设抽样概率仅取决于网络度数。
  • 开发一种方法,从通常被视为缺失或未观测到的RDS数据中提取隐藏社交网络的结构信息。
  • 将招募过程建模为包含招募事件时间序列、代金券使用和网络度数的连续时间随机过程。
  • 基于指数随机图模型,采用计算高效的框架估计隐藏网络结构。
  • 验证该方法在模型误设情况下的鲁棒性,特别是招募事件等待时间分布的误设。

提出的方法

  • 为RDS招募构建一个连续时间马尔可夫过程,其中隐藏网络中的每条边都有独立的指数分布等待时间以完成招募。
  • 引入兼容性条件(定义4),基于观测到的招募时间、度数和代金券使用情况,对招募诱导子图的可能结构施加约束。
  • 将招募诱导子图建模为指数随机图,其中边的概率取决于速率参数λ和可感染边的数量。
  • 采用贝叶斯框架,对λ使用伽马先验分布,其参数由基于观测数据推导出的边界确定,以确保先验的合理性。
  • 应用马尔可夫链蒙特卡洛中的Metropolis-Hastings算法,从隐藏图的后部分布中抽样,以实现网络结构的推断。
  • 采用经验贝叶斯方法,利用真实数据中的观测招募模式和度数分布,校准λ的先验分布。

实验结果

研究问题

  • RQ1RDS中的招募过程能否被建模为一个连续时间随机过程,以捕捉时间动态和网络结构?
  • RQ2在仅部分观测招募时间、度数和代金券使用的情况下,隐藏社交网络结构能在多大程度上被恢复?
  • RQ3当招募事件的等待时间分布假设错误时,该网络推断方法的鲁棒性如何?
  • RQ4所提出的模型是否通过整合网络拓扑和时间动态,优于现有RDS估计器?
  • RQ5该方法能否在真实RDS数据中准确估计边级招募速率λ并重建真实的网络拓扑?

主要发现

  • 该方法在正确设定指数等待时间时,能以高精度成功重建隐藏网络结构,平均准确率达到0.974。
  • 在模型正确设定下,真正例率(TPR)和真负例率(TNR)均保持在高水平(分别为0.974和0.990),表明边检测性能优异。
  • 在等待时间分布误设的情况下(如δ < 1或δ > 1的伽马分布),该方法在子图重建中仍保持鲁棒性,但λ估计存在偏差:δ < 1时高估,δ > 1时低估。
  • 在模型误设下,估计的边级招募速率λ存在偏差,平均估计值与真实λ = 1的偏差最大达0.10(δ = 0.1时)和0.01(δ = 0.01时)。
  • 基于圣彼得堡数据,经验贝叶斯对λ的先验分布被限定在9.8×10⁻⁴至4.2×10⁻²之间,确保了合理且信息丰富的先验分布。
  • 该方法在模拟数据和俄罗斯圣彼得堡注射吸毒者研究的真实RDS数据中均表现出色,具有高重建准确率和λ的稳定估计。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。