Skip to main content
QUICK REVIEW

[论文解读] Combining Experimental and Observational Data for Identification and Estimation of Long-Term Causal Effects

AmirEmad Ghassami, Liu, Chang|arXiv (Cornell University)|Jan 26, 2022
Advanced Causal Inference Techniques被引用 7
一句话总结

本文提出了三种新颖的数据融合方法,通过结合短期实验数据与长期观察数据来识别和估计长期因果效应,其中观察数据受到未观测混杂因素的影响。这些方法分别基于等混杂偏差、定制化工具变量或近似因果推断,每种方法均通过基于影响函数的估计量实现识别与估计,并具备双重稳健性特征。

ABSTRACT

We study identifying and estimating the causal effect of a treatment variable on a long-term outcome using data from an observational and an experimental domain. The observational data are subject to unobserved confounding. Furthermore, subjects in the experiment are only followed for a short period; thus, long-term effects are unobserved, though short-term effects are available. Consequently, neither data source alone suffices for causal inference on the long-term outcome, necessitating a principled fusion of the two. We propose three approaches for data fusion for the purpose of identifying and estimating the causal effect. The first assumes equal confounding bias for short-term and long-term outcomes. The second weakens this assumption by leveraging an observed confounder for which the short-term and long-term potential outcomes share the same partial additive association with this confounder. The third approach employs proxy variables of the latent confounder of the treatment-outcome relationship, extending the proximal causal inference framework to the data fusion setting. For each approach, we develop influence function-based estimators and analyze their robustness properties. We illustrate our methods by estimating the effect of class size on 8th-grade SAT scores using data from the Project STAR experiment combined with observational data from the Early Childhood Longitudinal Study.

研究动机与目标

  • 解决当实验数据仅捕捉短期结果而观察数据受未测量变量混杂影响时,估计长期因果效应的挑战。
  • 提出替代Athey等人(2020)所依赖的潜在无混杂性假设的识别策略,采用更灵活的假设。
  • 利用实验与观察数据的结合,实现对平均处理效应(ATE)和处理对接受者的效应(ETT)的估计。
  • 通过基于影响函数的估计量确保稳健性,该估计量在关键条件矩正确设定时具备双重稳健性。

提出的方法

  • 提出等混杂偏差假设,假设在加法或分位数-分位数形式下,短期与长期结果的混杂偏差程度相同。
  • 通过引入一种‘定制化工具变量’(BSIV)放松等混杂偏差假设,该变量为可观测的混杂因素,且对短期与长期结果具有相等的部分加法关联。
  • 应用近似因果推断框架,通过引入潜在国内混杂因素的代理变量,实现在较弱假设下的识别。
  • 为每种框架开发基于影响函数的估计量,确保当结果回归模型或倾向得分模型任一正确设定时,估计量具备双重稳健性。
  • 推导出结合两个数据域的估计方程,使用逆概率加权与双重稳健估计方程,以目标为ATE与ETT。
  • 在存在未观测混杂的结构因果模型下建立识别,基于条件独立性与矩条件假设。

实验结果

研究问题

  • RQ1当实验数据仅测量短期结果而观察数据受混杂影响时,能否识别长期因果效应?
  • RQ2如何利用等混杂偏差假设将数据融合中的短期与长期结果联系起来?
  • RQ3定制化工具变量(BSIV)能否在放松强等混杂偏差假设的同时仍实现识别?
  • RQ4通过潜在国内混杂因素的代理变量实现近似因果推断,能否在更弱假设下实现识别?
  • RQ5基于影响函数的估计量在该数据融合设置下的有限样本性质与稳健性特征为何?

主要发现

  • 所提出的等混杂偏差假设通过假设短期与长期结果的混杂偏差程度相等,实现了对长期因果效应的识别。
  • 定制化工具变量(BSIV)方法通过利用对两种结果具有相等部分关联的可观测变量,实现了对等混杂假设的放松。
  • 近似因果推断框架通过引入潜在国内混杂因素的代理变量,实现了在较弱结构假设下的识别,提供了一种灵活的替代方案。
  • 基于影响函数的估计量具备双重稳健性,当结果回归模型或倾向得分模型任一正确设定时,均能保持一致性。
  • 理论推导证实,在模型设定正确时,所提出的估计方程是无偏的,其一致性通过迭代期望定律得以建立。
  • 本文表明,即使单一数据源不足以提供有效估计,结合实验与观察数据仍可获得有效的长期因果估计。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。