Skip to main content
QUICK REVIEW

[论文解读] Adjusted Logistic Propensity Weighting Methods for Population Inference using Nonprobability Volunteer-Based Epidemiologic Cohorts

Lingxiao Wang, Richard Valliant|arXiv (Cornell University)|Jul 6, 2020
Statistical Methods and Bayesian Inference参考文献 17被引用 4
一句话总结

本文提出了一种调整逻辑倾向权重(ALP)方法,通过使用调整权重的逻辑回归估计参与率,实现从非概率、自愿参与的流行病学队列中进行总体水平推断。该方法在估计偏差和效率方面优于现有的逆倾向得分加权(IPSW)技术,采用泰勒线性化方差估计方法,全面考虑了所有变异来源,并在模拟研究和使用NHANES III与NHIS数据的真实数据应用中表现出稳健性能。

ABSTRACT

Many epidemiologic studies forgo probability sampling and turn to nonprobability volunteer-based samples because of cost, response burden, and invasiveness of biological samples. However, finite population inference is difficult to make from the nonprobability samples due to the lack of population representativeness. Aiming for making inferences at the population level using nonprobability samples, various inverse propensity score weighting (IPSW) methods have been studied with the propensity defined by the participation rate of population units in the nonprobability sample. In this paper, we propose an adjusted logistic propensity weighting (ALP) method to estimate the participation rates for nonprobability sample units. Compared to existing IPSW methods, the proposed ALP method is easy to implement by ready-to-use software while producing approximately unbiased estimators for population quantities regardless of the nonprobability sample rate. The efficiency of the ALP estimator can be further improved by scaling the survey sample weights in propensity estimation. Taylor linearization variance estimators are proposed for ALP estimators of finite population means that account for all sources of variability. The proposed ALP methods are evaluated numerically via simulation studies and empirically using the naïve unweighted National Health and Nutrition Examination Survey III sample, while taking the 1997 National Health Interview Survey as the reference, to estimate the 15-year mortality rates.

研究动机与目标

  • 解决由于代表性不足,导致无法从非概率、自愿参与的流行病学队列中进行有效有限总体推断的挑战。
  • 开发一种实用且高效的加权方法,在概率抽样不可行时,减少总体估计的偏差。
  • 通过引入调整权重和标准化,改进现有逆倾向得分加权(IPSW)方法,以提高估计器效率。
  • 提供一个方差估计框架,全面考虑ALP估计器中有限总体均值的所有变异来源。

提出的方法

  • 提出一种调整逻辑倾向权重(ALP)方法,通过辅助协变量的逻辑回归建模参与概率。
  • 在倾向模型中调整调查样本权重,以提高估计器效率并减少偏差。
  • 应用泰勒线性化方法推导ALP估计器的方差估计量,同时考虑抽样和加权的变异。
  • 采用两阶段方法:首先通过带缩放权重的逻辑回归估计倾向得分,然后对总体估计器应用逆概率加权。
  • 使用现成的软件进行实现,确保在现实流行病学研究中的实际可用性。
  • 通过模拟研究和基于NHANES III作为非概率样本、NHIS 1997作为参考总体的真实数据分析验证该方法。

实验结果

研究问题

  • RQ1当使用非概率自愿参与样本时,调整逻辑倾向权重能否产生对总体均值的近似无偏估计?
  • RQ2在不同非概率样本率下,ALP估计器的效率与现有IPSW方法相比如何?
  • RQ3在倾向模型中对调查权重进行缩放在多大程度上能提高估计器的精确度并减少偏差?
  • RQ4泰勒线性化方差估计量在多大程度上能准确捕捉ALP估计中全部的抽样变异?
  • RQ5ALP方法能否可靠地利用现实世界中的非概率数据估计总体水平结果(如15年死亡率)?

主要发现

  • 无论非概率样本率如何,ALP方法均能产生对总体均值的近似无偏估计,优于标准IPSW方法。
  • 在倾向模型中对调查样本权重进行缩放,显著提高了ALP估计器的效率。
  • ALP估计器的泰勒线性化方差估计量能有效涵盖所有变异来源,包括抽样和加权不确定性。
  • 模拟研究证实了ALP在不同数据生成机制和样本规模下的稳健性。
  • 使用NHANES III和NHIS 1997的真实数据分析表明,ALP能可靠估计15年死亡率,验证了其在现实世界中的实用性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。