[论文解读] Modeling and Analysing Respondent Driven Sampling as a Counting Process
本文提出了一种基于模型的响应者驱动抽样(RDS)推断框架,将招募过程视为连续时间计数过程,利用时间数据估计总体规模、度频次和患病率。通过使用强度函数对招募动态进行建模并借助最大似然估计,该方法在逆度加权的基础上显式建模了抽样偏差,实现了具有一致性和渐近正态性的估计量,并在有限样本中表现出更优性能。
Respondent-driven sampling (RDS) is an approach to sampling design and analysis which utilizes the networks of social relationships that connect members of the target population, using chain-referral methods to facilitate sampling. RDS typically leads to biased sampling, favoring participants with many acquaintances. Naive estimates, such as the sample average, which are uncorrected for the sampling bias, will themselves be biased. To compensate for this bias, current methodology suggests inverse-degree weighting, where the "degree" is the number of acquaintances. This stems from the fundamental RDS assumption that the probability of sampling an individual is proportional to their degree. Since this assumption is tenuous at best, we propose to harness the additional information encapsulated in the time of recruitment, into a model-based inference framework for RDS. This information is typically collected by researchers, but ignored. We adapt methods developed for inference in epidemic processes to estimate the population size, degree counts and frequencies. While providing valuable information in themselves, these quantities ultimately serve to debias other estimators, such a disease's prevalence. A fundamental advantage of our approach is that, being model-based, it makes all assumptions of the data-generating process explicit. This enables verification of the assumptions, maximum likelihood estimation, extension with covariates, and model selection. We develop asymptotic theory, proving consistency and asymptotic normality properties. We further compare these estimators to the standard inverse-degree weighting through simulations, and using real-world data. In both cases we find our estimators to outperform current methods. The likelihood problem in the model we present is convex, and thus efficiently solvable. We implement these estimators in an R package, chords, available on CRAN.
研究动机与目标
- 解决RDS的根本局限:对高度连接个体的抽样偏差。
- 通过利用未被充分利用的招募时间数据,克服标准RDS中度频次不可识别的问题。
- 开发一种基于模型的推断框架,明确表达所有假设,支持统计检验、估计与模型选择。
- 通过使用连续时间招募动态估计总体参数,改进逆度加权方法。
- 为患病率和度分布提供一致且渐近正态的估计量,有限样本性能更优。
提出的方法
- 将RDS招募建模为依赖于总体规模和特定度招募率的连续时间计数过程,使用强度函数表示。
- 基于计数过程理论(Andersen et al., 1995),推导出观测招募序列及其时间的似然函数。
- 通过最大似然估计度频次 $ f_k $ 和患病率 $ p_k $,利用从对数似然导数导出的根查找方程求解 $ \hat{N}_k $。
- 应用delta方法推导估计患病率 $ \widehat{H} = \sum_k \hat{f}_k \hat{p}_k $ 的渐近正态性,确保有效的统计推断。
- 推导出坐标方向凸的似然问题,实现高效且全局收敛的优化。
- 在R包 'chords' 中实现该方法,该包已发布于CRAN,便于实际应用。
实验结果
研究问题
- RQ1能否利用招募时间数据识别RDS中此前无法识别的参数(如度频次)?
- RQ2将RDS建模为计数过程是否能比标准的逆度加权方法提供更准确、更少偏差的总体患病率估计?
- RQ3具有明确假设的基于模型的框架是否能相比临时加权方案提高RDS推断的可靠性和有效性?
- RQ4在正则条件下,所提出的基于似然的估计量是否具有一致性和渐近正态性?
- RQ5在模拟和真实数据中,新估计量在有限样本中的表现与现有方法相比如何?
主要发现
- 在模型设定良好的前提下,所提出的方法能实现总体患病率和度频次的一致且渐近正态的估计。
- 似然函数具有坐标方向凸性,确保全局收敛性与最大似然估计的高效计算。
- 模拟研究显示,新估计量在偏差和均方误差方面优于标准的逆度加权方法。
- 真实数据分析表明,新方法能产生更可靠的患病率估计,且具有更好的覆盖性能。
- 该方法支持模型选择、假设检验以及协变量的引入,提供了一个灵活且透明的推断框架。
- R包 'chords' 提供了该方法的实际实现,使应用研究者能够便捷使用。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。