Skip to main content
QUICK REVIEW

[论文解读] Data-Driven Causal Effect Estimation Based on Graphical Causal Modelling: A Survey

Debo Cheng, Jiuyong Li|arXiv (Cornell University)|Aug 20, 2022
Advanced Causal Inference Techniques被引用 4
一句话总结

本综述系统性地回顾了基于图形因果模型的数据驱动因果效应估计方法,重点关注在部分或不确定的因果知识下对平均处理效应(ATE)的估计。它整合了核心理论,按方法对混杂变量和潜变量的处理方式进行分类,并评估了其假设、优势与局限性,为从观测数据中进行因果推断的未来研究提供了基础。

ABSTRACT

In many fields of scientific research and real-world applications, unbiased estimation of causal effects from non-experimental data is crucial for understanding the mechanism underlying the data and for decision-making on effective responses or interventions. A great deal of research has been conducted to address this challenging problem from different angles. For estimating causal effect in observational data, assumptions such as Markov condition, faithfulness and causal sufficiency are always made. Under the assumptions, full knowledge such as, a set of covariates or an underlying causal graph, is typically required. A practical challenge is that in many applications, no such full knowledge or only some partial knowledge is available. In recent years, research has emerged to use search strategies based on graphical causal modelling to discover useful knowledge from data for causal effect estimation, with some mild assumptions, and has shown promise in tackling the practical challenge. In this survey, we review these data-driven methods on causal effect estimation for a single treatment with a single outcome of interest and focus on the challenges faced by data-driven causal effect estimation. We concisely summarise the basic concepts and theories that are essential for data-driven causal effect estimation using graphical causal modelling but are scattered around the literature. We identify and discuss the challenges faced by data-driven causal effect estimation and characterise the existing methods by their assumptions and the approaches to tackling the challenges. We analyse the strengths and limitations of the different types of methods and present an empirical evaluation to support the discussions. We hope this review will motivate more researchers to design better data-driven methods based on graphical causal modelling for the challenging problem of causal effect estimation.

研究动机与目标

  • 解决在缺乏完整因果知识(如完整因果图或所有混杂因子)时,从观测数据中 unbiased 地估计因果效应的挑战。
  • 整合与数据驱动因果效应估计相关的图形因果建模的零散理论基础,特别是基于忠实性(faithfulness)和马尔可夫条件(Markov condition)等假设的理论。
  • 识别并分析数据驱动因果效应估计中的关键挑战,包括因果结构学习中的不确定性、计算复杂性以及潜混杂偏倚。
  • 基于其假设、方法论方法和性能,评估现有方法,重点关注平均处理效应(ATE)的估计。
  • 提供结构化的概述与实证评估,以指导研究人员选择并开发更优的数据驱动因果推断方法。

提出的方法

  • 系统性地回顾图形因果模型(如DAGs、MAGs)作为在观测数据中表示因果关系的基础。
  • 根据方法对因果结构学习中不确定性的处理方式,对数据驱动方法进行分类,例如使用基于约束的方法、基于评分的方法或混合搜索策略。
  • 分析引入工具变量(IVs)和条件工具变量(conditional IVs)的方法,以应对潜混杂偏倚,包括可容忍无效工具的策略(如sisVIVE)。
  • 利用具有已知真实因果效应和结构的合成数据集和半合成数据集(如IHDP、Twins)评估方法,以检验其准确性和鲁棒性。
  • 评估方法的时间复杂度和可扩展性,特别是在存在潜变量的高维设置下。
  • 通过真实世界数据集(如Job training、401(k)、Schoolingreturns)的实证评估比较方法,其中真实因果效应通过领域知识近似获得。

实验结果

研究问题

  • RQ1当完整因果结构未知或部分观测时,基于图形因果建模的数据驱动方法如何估计平均处理效应(ATE)?
  • RQ2这些方法所依赖的关键假设(如忠实性、马尔可夫性、因果充分性)是什么,它们如何影响估计的准确性?
  • RQ3在从观测数据中进行因果效应估计时,这些方法如何处理潜混杂因子和无效工具变量?
  • RQ4在数据驱动因果推断中,方法论复杂性、计算效率与估计准确性之间的权衡是什么?
  • RQ5在真实世界数据集中缺乏真实因果效应的情况下,如何有意义地开展实证评估?

主要发现

  • 基于图形因果建模的数据驱动因果效应估计方法即使在因果结构知识不完整的情况下,也能有效估计ATE,尤其当与工具变量技术结合时效果更佳。
  • 如sisVIVE等方法对无效工具的存在表现出鲁棒性,使其适用于存在未测量混杂因子的实际应用场景。
  • 在IHDP和Twins等半合成数据集上的实证评估表明,许多数据驱动方法能达到合理的准确性,但性能在很大程度上取决于数据质量和底层假设。
  • 真实世界数据集中缺乏真实因果效应限制了可靠评估,现有基准依赖于领域知识或经验估计作为代理。
  • 时间复杂度仍是重大挑战,特别是在高维设置下的结构学习中,限制了部分方法的可扩展性。
  • 本综述指出,当前方法仍受忠实性与因果充分性等假设的制约,未来研究需在保持估计有效性的前提下进一步放宽这些假设。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。