[论文解读] On the Use of Auxiliary Variables in Multilevel Regression and Poststratification
本文提出了一种集成的MRP框架,通过在辅助变量上条件化建模纳入概率,并将估计的纳入概率作为结果模型中的预测变量,从而提升多层回归与后分层(MRP)的性能。该方法在模型误设情况下显著提高了估计精度,并在ABCD研究数据中展示了稳健的有限总体推断性能,其中MRP-INT在小样本子群体中表现出更优的偏差控制能力,优于标准MRP。
Multilevel regression and poststratification (MRP) is a popular method for addressing selection bias in subgroup estimation, with broad applications across fields from social sciences to public health. In this paper, we examine the inferential validity of MRP in finite populations, exploring the impact of poststratification and model specification. The success of MRP relies heavily on the availability of auxiliary information that is strongly related to the outcome. To enhance the fitting performance of the outcome model, we recommend modeling the inclusion probabilities conditionally on auxiliary variables and incorporating flexible functions of estimated inclusion probabilities as predictors in the mean structure. We present a statistical data integration framework that offers robust inferences for probability and nonprobability surveys, addressing various challenges in practical applications. Our simulation studies indicate the statistical validity of MRP, which involves a tradeoff between bias and variance, with greater benefits for subgroup estimates with small sample sizes, compared to alternative methods. We have applied our methods to the Adolescent Brain Cognitive Development (ABCD) Study, which collected information on children across 21 geographic locations in the U.S. to provide national representation, but is subject to selection bias as a nonprobability sample. We focus on the cognition measure of diverse groups of children in the ABCD study and show that the use of auxiliary variables affects the findings on cognitive performance.
研究动机与目标
- 为通过改进有限总体中多层回归与后分层(MRP)的推断有效性,解决非概率调查中的选择偏差。
- 研究辅助变量选择与模型设定对MRP性能的影响,特别是在样本量有限的小样本子群体中。
- 开发一种数据整合框架,利用基于模型的推断方法,结合概率与非概率调查数据。
- 评估在MRP结果模型中包含估计的纳入概率作为预测变量的有效性,以在模型误设下减少偏差。
- 为公共卫生与社会科学实证研究在现实世界应用中提供辅助变量建模的实用指导。
提出的方法
- 使用逻辑回归或类似模型,基于辅助变量条件化建模纳入概率,以估计选择倾向。
- 将估计的纳入概率作为结果模型中的预测变量,通过在纳入概率的离散取值上设置随机截距,实现灵活建模。
- 采用具有分层先验的多层回归框架,对估计进行正则化,提升小区域预测能力。
- 使用来自参考概率样本(如ACS)的已知总体单元格大小进行后分层,将样本估计校准至目标总体。
- 实施合成总体生成方法(如WFPBB算法),结合多重插补以估计方差并考虑不确定性。
- 使用合并规则(如Dong等,2014)对多个合成总体的结果进行聚合,并计算稳健的置信区间。
实验结果
研究问题
- RQ1在模型误设条件下,将估计的纳入概率作为结果模型中的预测变量,对MRP性能有何影响?
- RQ2辅助变量选择与模型设定对小样本子群体中MRP估计的偏差与方差有何影响?
- RQ3所提出的集成MRP(MRP-INT)框架与标准MRP及逆概率加权法相比,在估计精度与稳健性方面表现如何?
- RQ4包含能预测选择的辅助变量在多大程度上可改善非概率样本中的有限总体推断?
- RQ5当结果模型被误设时,所提出的框架能否在代表性不足的少数族裔群体中产生有效且稳健的推断?
主要发现
- MRP-INT(将估计的纳入概率作为预测变量)在小样本子群体中产生的估计更加稳定且偏差更小,优于MRP-P与MRP-R,尤其在样本量较小的子群体中表现更优。
- 对于贫困白人儿童子群体(n=291),MRP-INT在估计的认知得分上与完整模型基线的差异小于MRP-P,表明在模型误设下具有更高的稳健性。
- 对于来自大家庭(家庭规模>5)且父母未在劳动力市场中的儿童(n=132),即使从结果模型中移除收入与劳动力市场状况,MRP-INT的估计仍保持一致;而MRP-P的估计则发生显著变化,表明其对模型误设具有更强的保护能力。
- MRP-INT的总体认知得分估计为85.91(95%置信区间:85.75, 86.07),与MRP-P和MRP结果接近,但在子群体估计中表现出更高的精度与稳健性。
- 逆概率加权(IPW)方法得到的估计值更高(86.20),置信区间更宽(86.03, 86.36),表明其方差更大,可能存在不稳定性,相较基于MRP的方法表现较差。
- 模拟研究与ABCD研究结果共同表明,MRP-INT通过不仅用于后分层,还用于改进结果模型预测,有效降低了子群体估计的偏差。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。