[论文解读] Collaborative-controlled LASSO for Constructing Propensity Score-based Estimators in High-Dimensional Data
本文提出一种协作控制的LASSO方法用于高维数据中的倾向得分估计,采用C-TMLE联合优化治疗预测与因果效应估计。结果表明,基于C-TMLE的模型选择在减少平均治疗效应估计的偏差和提高置信区间覆盖率方面优于传统的外部交叉验证。
Propensity score (PS) based estimators are increasingly used for causal inference in observational studies. However, model selection for PS estimation in high-dimensional data has received little attention. In these settings, PS models have traditionally been selected based on the goodness-of-fit for the treatment mechanism itself, without consideration of the causal parameter of interest. Collaborative minimum loss-based estimation (C-TMLE) is a novel methodology for causal inference that takes into account information on the causal parameter of interest when selecting a PS model. This "collaborative learning" considers variable associations with both treatment and outcome when selecting a PS model in order to minimize a bias-variance trade off in the estimated treatment effect. In this study, we introduce a novel approach for collaborative model selection when using the LASSO estimator for PS estimation in high-dimensional covariate settings. To demonstrate the importance of selecting the PS model collaboratively, we designed quasi-experiments based on a real electronic healthcare database, where only the potential outcomes were manually generated, and the treatment and baseline covariates remained unchanged. Results showed that the C-TMLE algorithm outperformed other competing estimators for both point estimation and confidence interval coverage. In addition, the PS model selected by C-TMLE could be applied to other PS-based estimators, which also resulted in substantive improvement for both point estimation and confidence interval coverage. We illustrate the discussed concepts through an empirical example comparing the effects of non-selective nonsteroidal anti-inflammatory drugs with selective COX-2 inhibitors on gastrointestinal complications in a population of Medicare beneficiaries.
研究动机与目标
- 为解决外部交叉验证在高维数据倾向得分模型选择中的局限性。
- 通过将目标因果参数整合到模型选择中,而非仅关注治疗预测,以改进因果推断。
- 评估基于C-TMLE的LASSO模型选择是否在高维设置下提升点估计和置信区间表现。
- 评估协作选择模型向其他基于倾向得分的估计器的可转移性。
- 通过真实电子医疗健康数据和准实验模拟验证研究结果。
提出的方法
- 提出一种协作控制LASSO框架,利用C-TMLE选择高维倾向得分模型中LASSO的调参。
- 采用基于目标损失的估计(TMLE)并结合协作学习,以在治疗效应估计中平衡偏差与方差。
- 使用外部交叉验证进行初始模型拟合,但应用C-TMLE根据感兴趣的因果参数优化模型选择。
- 在C-TMLE算法中采用两种初始估计器:一个朴素模型和一个包含基线及高维倾向得分(hdPS)协变量的超学习模型。
- 通过交叉验证的二项偏差评估模型表现,并比较不同估计器在估计精度和置信区间覆盖率方面的表现。
- 将所选倾向得分模型应用于多种基于倾向得分的估计器(如IPW、AIPW),以检验协作选择的泛化能力。
实验结果
研究问题
- RQ1在高维设置下,通过C-TMLE进行协作模型选择是否能改善平均治疗效应的点估计,相比外部交叉验证?
- RQ2基于C-TMLE的LASSO模型选择如何影响因果效应估计器的置信区间覆盖率和长度?
- RQ3通过C-TMLE选择的倾向得分模型能否有效转移至非协作倾向得分估计器,以提升其性能?
- RQ4不同的初始估计器(朴素模型 vs. 超学习模型)如何影响C-TMLE中的最终模型选择和估计精度?
- RQ5当因果推断的目标是混杂控制时,外部交叉验证是否在倾向得分模型选择中表现次优?
主要发现
- 在模拟中,C-TMLE1和C-TMLE0估计器在点估计精度和置信区间覆盖率方面均表现最佳。
- C-TMLE选择的模型包含166个协变量,λ = 0.000238,选择的协变量多于CV.LASSO,尤其包括预测能力较弱但具有强混杂效应的hdPS变量。
- 当应用于其他基于倾向得分的估计器时,协作选择的模型显著改善了点估计,但对原始IPW估计器例外。
- 使用朴素初始估计器的C-TMLE1估计器的交叉验证二项偏差为1.199632,略高于CV.LASSO的1.199288,但因果推断表现更优。
- 实证分析显示,与非选择性NSAIDs相比,COX-2抑制剂的平均加法治疗效应估计为-0.249%,但不具有统计学显著性。
- 研究结论认为,外部交叉验证不足以实现最优混杂控制,基于C-TMLE的集成学习为倾向得分模型选择提供了更优替代方案。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。