[论文解读] Network Structure and Biased Variance Estimation in Respondent Driven Sampling
本文指出,当潜在社交网络违反一阶马尔可夫(FOM)假设时,响应者驱动抽样(RDS)的方差估计器系统性地低估了抽样方差。基于来自 Facebook 和 Add Health 的实证网络数据,作者证明所有 215 个分析网络均违反 FOM 假设,导致方差估计偏差,并提出了两种替代估计器,尽管这些估计器可减少但无法完全消除该偏差,原因在于网络信息有限。
This paper explores bias in the estimation of sampling variance in Respondent Driven Sampling (RDS). Prior methodological work on RDS has focused on its problematic assumptions and the biases and inefficiencies of its estimators of the population mean. Nonetheless, researchers have given only slight attention to the topic of estimating sampling variance in RDS, despite the importance of variance estimation for the construction of confidence intervals and hypothesis tests. In this paper, we show that the estimators of RDS sampling variance rely on a critical assumption that the network is First Order Markov (FOM) with respect to the dependent variable of interest. We demonstrate, through intuitive examples, mathematical generalizations, and computational experiments that current RDS variance estimators will always underestimate the population sampling variance of RDS in empirical networks that do not conform to the FOM assumption. Analysis of 215 observed university and school networks from Facebook and Add Health indicates that the FOM assumption is violated in every empirical network we analyze, and that these violations lead to substantially biased RDS estimators of sampling variance. We propose and test two alternative variance estimators that show some promise for reducing biases, but which also illustrate the limits of estimating sampling variance with only partial information on the underlying population social network.
研究动机与目标
- 调查网络结构对方差估计在响应者驱动抽样(RDS)中影响。
- 识别当前 RDS 方差估计器所依赖的关键假设:一阶马尔可夫(FOM)特性。
- 评估现实世界社交网络在多大程度上违反 FOM 假设,从而影响方差估计。
- 提出并评估在非 FOM 网络结构下可减少偏差的替代方差估计器。
- 强调当仅掌握部分网络信息时,RDS 中方差估计的局限性。
提出的方法
- 理论分析表明,RDS 方差估计器依赖于 FOM 假设,即节点被招募的概率仅取决于其在招募链中的直接前驱。
- 通过数学推广和计算实验,表明 FOM 假设的违反会导致抽样方差的低估。
- 基于来自 Facebook 和 Add Health 的 215 个现实世界社交网络进行实证验证,以检验 FOM 假设和方差估计偏差。
- 基于网络结构校正,提出两种新的方差估计器,整合了招募过程中的高阶依赖关系。
- 通过模拟和与标准 RDS 方差估计器的比较,评估新估计器的性能。
- 研究使用网络度量和统计建模,量化在不同网络拓扑下方差估计的偏差。
实验结果
研究问题
- RQ1在 RDS 研究中使用的现实世界社交网络中,一阶马尔可夫(FOM)假设是否成立?
- RQ2违反 FOM 假设在多大程度上导致 RDS 中抽样方差估计的偏差?
- RQ3当 FOM 假设被违反时,能否通过替代方差估计器减少偏差?
- RQ4网络结构如聚类和度异质性如何影响 RDS 方差估计?
- RQ5当仅掌握底层网络的部分知识时,RDS 中方差估计的实用改进极限是什么?
主要发现
- 从 Facebook 和 Add Health 分析的全部 215 个实证网络均违反一阶马尔可夫(FOM)假设,表明标准 RDS 方差估计器的基础条件普遍失效。
- 在不符合 FOM 假设的网络中,当前 RDS 方差估计器始终低估真实抽样方差。
- 低估的幅度显著,偏差水平因网络结构而异,但在所有分析案例中均持续存在。
- 所提出的替代方差估计器相比标准 RDS 估计器减少了偏差,但由于网络信息不完整,无法完全消除偏差。
- 本研究表明,根本挑战在于仅掌握部分网络数据时估计抽样方差,凸显了 RDS 方法论中的固有局限性。
- 结果强调了在构建置信区间或进行假设检验时,对 RDS 研究采取方法论谨慎态度的必要性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。