Skip to main content
QUICK REVIEW

[论文解读] Semiparametric Proximal Causal Inference

Yifan Cui, Hongming Pu|arXiv (Cornell University)|Nov 17, 2020
Advanced Causal Inference Techniques参考文献 41被引用 18
一句话总结

本文为近似因果推断发展了半参数理论,通过利用处理和结果的代理变量,在未观测混杂因素下实现了平均处理效应的识别与高效估计。研究建立了非参数识别、效率边界,并提出了双重稳健估计量,在模拟和一项关于右心导管检查的ICU真实世界研究中,其表现优于标准方法。

ABSTRACT

Skepticism about the assumption of no unmeasured confounding, also known as exchangeability, is often warranted in making causal inferences from observational data; because exchangeability hinges on an investigator's ability to accurately measure covariates that capture all potential sources of confounding. In practice, the most one can hope for is that covariate measurements are at best proxies of the true underlying confounding mechanism operating in a given observational study. In this paper, we consider the framework of proximal causal inference introduced by Miao et al. (2018); Tchetgen Tchetgen et al. (2020), which while explicitly acknowledging covariate measurements as imperfect proxies of confounding mechanisms, offers an opportunity to learn about causal effects in settings where exchangeability on the basis of measured covariates fails. We make a number of contributions to proximal inference including (i) an alternative set of conditions for nonparametric proximal identification of the average treatment effect; (ii) general semiparametric theory for proximal estimation of the average treatment effect including efficiency bounds for key semiparametric models of interest; (iii) a characterization of proximal doubly robust and locally efficient estimators of the average treatment effect. Moreover, we provide analogous identification and efficiency results for the average treatment effect on the treated. Our approach is illustrated via simulation studies and a data application on evaluating the effectiveness of right heart catheterization in the intensive care unit of critically ill patients.

研究动机与目标

  • 解决观测研究中测量协变量仅为真实混杂因素代理变量时的未观测混杂问题。
  • 开发一个通用的半参数框架,用于在近似识别假设下识别和估计平均处理效应。
  • 推导近似因果推断的效率边界,并构建双重稳健、局部高效估计量。
  • 将该框架扩展至处理组上的平均处理效应(ATT),以扩大其在政策相关效应量中的适用性。

提出的方法

  • 提出了一组新的非参数条件,用于利用处理和结果代理变量实现平均处理效应的近似识别。
  • 发展了通用的半参数理论,包括关键模型的高效影响函数和渐近方差边界。
  • 刻画了近似双重稳健和局部高效估计量的性质,其一致性在结果模型或倾向得分模型任一正确时均能保持。
  • 提出了一种新颖的近似逆概率加权(PIPW)估计量和一种近似双重稳健(PDR)估计量,以提升稳健性。
  • 通过模拟研究评估了在各种模型误设情形和不同代理变量强度下的性能表现。
  • 将该方法应用于一项关于危重患者右心导管检查的真实数据集,评估了代理变量选择的敏感性。

实验结果

研究问题

  • RQ1在弱于标准可交换性假设的条件下,近似因果推断能否实现非参数识别?
  • RQ2近似估计量的平均处理效应的半参数效率边界是什么?
  • RQ3如何构建近似双重稳健和局部高效估计量?其有限样本性质如何?
  • RQ4在模型误设条件下,近似框架是否相比标准双重稳健方法能提升估计精度和覆盖性能?
  • RQ5在真实世界应用中,近似估计对代理变量选择的敏感性如何?

主要发现

  • 在模型误设的模拟情景中,近似估计量(尤其是PDR和POR)相较于标准双重稳健估计量,表现出更优的偏差和覆盖性能。
  • 在情景1中,近似PDR估计量实现了96.8%的覆盖区间,均方误差(MSE)为1.0,优于标准DR估计量(22.4%覆盖,MSE 0.6)。
  • 当代理变量强度降低时(σwz = 0.15),近似估计量仍具竞争力,PDR在情景4中保持88.8%的覆盖,而标准DR下降至20.2%。
  • 在真实数据应用中,当仅使用paco21作为代理变量时,近似估计值显著变化,表明其可能是弱代理或模型误设的混杂代理。
  • 在情景3中,近似IPW估计量得到正值但不显著的结果(0.41,95%置信区间:-0.07, 0.90),提示当代理变量较弱时可能存在模型误设。
  • 敏感性分析确认,代理变量选择对结果影响重大:以pafi1为代理时估计结果稳定,而使用paco21则导致发散且不可靠的估计。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。