Skip to main content
QUICK REVIEW

[论文解读] Respondent-Driven Sampling: An Assessment of Current Methodology

Krista J. Gile, Mark S. Handcock|ArXiv.org|Apr 12, 2009
HIV, Drug Use, Sexual Risk参考文献 20被引用 5
一句话总结

本文批判性地評估了回應者驅動抽樣(RDS)方法論,揭示其廣泛聲稱的漸近無偏性依賴於不切實際的假設。研究顯示,RDS估計量因抽樣波次不足、受訪者推薦行為,以及將無放回抽樣錯誤地近似為有放回隨機遊走而產生偏誤——尤其在樣本佔比較大時更為顯著。

ABSTRACT

Respondent-Driven Sampling (RDS) employs a variant of a link-tracing network sampling strategy to collect data from hard-to-reach populations. By tracing the links in the underlying social network, the process exploits the social structure to expand the sample and reduce its dependence on the initial (convenience) sample. The primary goal of RDS is typically to estimate population averages in the hard-to-reach population. The current estimates make strong assumptions in order to treat the data as a probability sample. In particular, we evaluate three critical sensitivities of the estimators: to bias induced by the initial sample, to uncontrollable features of respondent behavior, and to the without-replacement structure of sampling. This paper sounds a cautionary note for the users of RDS. While current RDS methodology is powerful and clever, the favorable statistical properties claimed for the current estimates are shown to be heavily dependent on often unrealistic assumptions.

研究动机与目标

  • 評估在現實抽樣限制下RDS估計量的統計有效性。
  • 研究初始便利抽樣對RDS估計量偏誤的影響。
  • 檢討未受控的受訪者行為(特別是偏好推薦)對估計精確度的影響。
  • 評估將無放回抽樣建模為有放回隨機遊走的不準確性在RDS中的影響。
  • 強調需改進受訪者行為與網絡結構的資料收集,以提升RDS的可靠性。

提出的方法

  • 使用模擬研究,評估在不同抽樣波次數量與群體聚集程度下RDS估計量的偏誤。
  • 分析受訪者偏好推薦行為對估計量偏誤的影響,特別是在推薦非隨機時的情況。
  • 比較有放回隨機遊走模型與真實無放回抽樣過程下RDS估計量的表現。
  • 採用馬爾可夫鏈表示法來建模抽樣過程中的波次混合行為。
  • 建議改進度數詢問與補充調查問題,以減少估計偏誤。
  • 回顧現有估計量,並識別其失效的條件,特別是在樣本佔比大且存在同質性時。

实验结果

研究问题

  • RQ1RDS中抽樣波次的數量在多大程度上能減少來自初始便利樣本的偏誤?
  • RQ2受訪者偏好推薦行為如何在RDS估計量中引入系統性偏誤?
  • RQ3將有放回隨機遊走近似應用於實際的無放回RDS抽樣時,其不準確性如何?
  • RQ4在哪些群體條件下(如高聚集性或活動水平差異)RDS估計量會表現出顯著偏誤?
  • RQ5需要哪些資料收集與建模改進,才能確保RDS推論的可靠性?

主要发现

  • RDS通常使用的抽樣波次數量不足以實現該方法聲稱的漸近無偏性,特別是在高度聚集的群體中。
  • RDS估計量對偏好推薦行為極為敏感,可能引入顯著偏誤,而現有模型未加以考慮。
  • 當目標群體中大規模樣本被抽樣時,將RDS近似為有放回隨機遊走會導致顯著偏誤,特別是在同質性與活動水平差異存在時。
  • 在樣本佔比大的情況下,樣本平均數往往優於現行RDS估計量,顯示估計量設計存在根本性缺陷。
  • 現行RDS推論框架依賴於不切實際的假設,其有利的統計性質在實際中並不能保證。
  • 改進受訪者行為(如度數詢問與推薦模式)的資料收集,對提升RDS估計的有效性與可靠性至關重要。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。