[论文解读] Estimating Hidden Population Size using Respondent-Driven Sampling Data
本文提出一种贝叶斯方法,仅使用受访者抽样(RDS)数据来估计隐藏人群规模,利用抽样过程中观察到的个人网络规模递减序列。该方法采用连续抽样近似来建模人群耗竭过程,得到具有良好频率覆盖度的可靠人群规模估计,并改善对总体特征的推断。
Respondent-Driven Sampling (RDS) is an approach to sampling design and inference in hard-to-reach human populations. Typically, a sampling frame is not available, and population members are difficult to identify or recruit from broader sampling frames. Common examples include injecting drug users, men who have sex with men, and female sex workers. Most analysis of RDS data has focused on estimating aggregate characteristics, such as disease prevalence. However, RDS is often conducted in settings where the population size is unknown and of great independent interest. This paper presents an approach to estimating the size of a target population based on data collected through RDS. The proposed approach uses a successive sampling approximation to RDS to leverage information in the ordered sequence of observed personal network sizes. The inference uses the Bayesian framework, allowing for the incorporation of prior knowledge. A flexible class of priors for the population size is proposed that aids elicitation. An extensive simulation study provides insight into the performance of the method for estimating population size under a broad range of conditions. A further study shows the approach also improves estimation of aggregate characteristics. A particular choice of the prior produces interval estimates with good frequentist properties. Finally, the method demonstrates sensible results when used to estimate the numbers of sub-populations most at risk for HIV in two cities in El Salvador.
研究动机与目标
- 开发一种仅使用RDS数据、无需额外数据源的硬性接触人群规模估计方法。
- 利用RDS中个人网络规模的有序序列推断人群规模,将抽样依赖性视为信息性而非干扰因素。
- 提供一个灵活的贝叶斯框架,整合先验知识,并支持与其他估计方法的连贯结合。
- 通过使用更准确的人群规模估计,改进对总体特征(如疾病患病率)的估计。
- 在多种抽样条件和真实应用场景下(包括萨尔瓦多的HIV高危人群),展示该方法的性能。
提出的方法
- 采用连续抽样近似建模RDS过程,其中较大的网络规模更可能在早期被抽中,反映人群耗竭现象。
- 应用贝叶斯分层模型估计人群规模N,采用灵活的先验类以辅助先验 elicitation 并提高稳健性。
- 将观察到的个人网络规模序列建模为抽样顺序的函数,假设规模递减表示人群耗竭。
- 使用马尔可夫链蒙特卡洛(MCMC)进行后验计算,实现对N的完整不确定性量化。
- 通过将该方法的后验分布作为先验,与其它数据源(如捕获-再捕获、乘数法)结合,实现集成推断。
- 通过在不同人群规模、网络结构和抽样比例下进行大量模拟研究验证该方法。
实验结果
研究问题
- RQ1能否仅从RDS数据中可靠估计人群规模,而无需额外数据源?
- RQ2在RDS下,观察到的个人网络规模的序列模式如何用于人群规模估计?
- RQ3所提出的贝叶斯方法在不同抽样条件下,其频率覆盖度和偏差表现如何?
- RQ4使用估计的人群规模能否提高基于RDS的患病率估计器的精度?
- RQ5该方法在真实世界应用中的表现如何,例如在萨尔瓦多估计HIV高危人群规模?
主要发现
- 该方法在具有强差异活动性的挑战性条件下,仍能产生具有良好频率覆盖度的人群规模区间估计。
- 模拟结果表明,该方法表现良好,尤其在样本比例较大及网络规模异质性较强时。
- 在SS估计量(Gile, 2011)中使用估计的人群规模,显著提高了患病率估计的精度。
- 该方法对萨尔瓦多HIV高危人群的估计结果与UNAIDS指南及捕获-再捕获估计结果相容。
- 该方法的后部分布可用作多方法推断中的先验,实现人群规模估计的连贯且逐步改进。
- 在信息有限的情况下,该方法产生较宽的可信区间,真实反映了不确定性,相比低估变异性的方法更具诚实性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。