[论文解读] On Adaptive Propensity Score Truncation in Causal Inference
本文提出 Positivity-C-TMLE,一种数据自适应的倾向得分截断方法,通过协作目标最大似然估计最小化目标参数估计量的均方误差,选择最优截断点。该方法在有限样本中优于固定截断和其他自适应方法,在实际正性性违反下实现了更优的点估计和置信区间覆盖。
The positivity assumption, or the experimental treatment assignment (ETA) assumption, is important for identifiability in causal inference. Even if the positivity assumption holds, practical violations of this assumption may jeopardize the finite sample performance of the causal estimator. One of the consequences of practical violations of the positivity assumption is extreme values in the estimated propensity score (PS). A common practice to address this issue is truncating the PS estimate when constructing PS-based estimators. In this study, we propose a novel adaptive truncation method, Positivity-C-TMLE, based on the collaborative targeted maximum likelihood estimation (C-TMLE) methodology. We demonstrate the outstanding performance of our novel approach in a variety of simulations by comparing it with other commonly studied estimators. Results show that by adaptively truncating the estimated PS with a more targeted objective function, the Positivity-C-TMLE estimator achieves the best performance for both point estimation and confidence interval coverage among all estimators considered.
研究动机与目标
- 解决由于实际正性性假设违反导致的因果推断中有限样本性能问题。
- 开发一种基于目标参数而非固定阈值的数据自适应截断方法,以选择最优倾向得分截断点。
- 在强实际正性性违反下提升因果估计量的稳健性与效率。
- 通过模拟研究评估所提方法相对于现有截断策略的性能。
- 为小样本中置信区间构建提供更可靠的方差估计量。
提出的方法
- 该方法将协作目标最大似然估计(C-TMLE)扩展至将最优倾向得分截断分位数 γ 作为调优参数进行选择。
- 将截断选择建模为一个以因果参数估计量均方误差为目标的优化问题。
- 算法采用基于目标损失的估计框架,迭代更新倾向得分与截断水平。
- 通过最小化直接反映最终因果估计偏差-方差权衡的损失函数,选择最优 γ。
- 该方法结合了双重稳健估计方程,并利用影响曲线进行方差估计。
- 应用交叉验证与模型验证评估性能,重点在于最小化有限样本中的估计误差。
实验结果
研究问题
- RQ1数据自适应的倾向得分截断方法在均方误差与置信区间覆盖方面是否优于固定截断策略?
- RQ2在实际正性性违反下,截断点的选择如何影响因果估计量的偏差与方差?
- RQ3基于目标参数估计量优化截断水平,是否能带来优于仅基于倾向得分模型优化的有限样本性能?
- RQ4所提估计量的方差估计与现有方法相比如何,特别是在强正性性违反的小样本中?
- RQ5自适应截断方法的性能是否对不同程度的实际正性性违反具有鲁棒性?
主要发现
- Positivity-C-TMLE 在所有模拟情景下均实现了最优的均方误差与置信区间覆盖表现。
- 该方法选择的截断点对正性参数 C 的敏感度高于交叉验证,而对样本量的敏感度较低。
- Positivity-C-TMLE 的方差估计量偏差小于 CV-TMLE 与 MV-TMLE,在 C=2 时,估计标准误与真实标准误的比值为 0.88。
- CV-TMLE 与 MV-TMLE 的方差低估更为严重,在强正性性违反下(C=2)比值分别降至 0.51 与 0.49。
- 在小样本中强正性性违反(N=200, C=2)时,Positivity-C-TMLE 在偏差与均方误差方面显著优于 MV-TMLE。
- 倾向得分的负对数似然被证明是选择截断点的不良目标函数,进一步证实了直接针对因果参数进行优化的价值。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。