[Paper Review] Data-Driven Causal Effect Estimation Based on Graphical Causal Modelling: A Survey
This survey presents a comprehensive review of data-driven causal effect estimation methods using graphical causal models, focusing on average treatment effect (ATE) estimation under partial or uncertain causal knowledge. It synthesizes core theories, categorizes methods by their handling of confounders and latent variables, and evaluates their assumptions, strengths, and limitations, offering a foundation for future research in causal inference from observational data.
In many fields of scientific research and real-world applications, unbiased estimation of causal effects from non-experimental data is crucial for understanding the mechanism underlying the data and for decision-making on effective responses or interventions. A great deal of research has been conducted to address this challenging problem from different angles. For estimating causal effect in observational data, assumptions such as Markov condition, faithfulness and causal sufficiency are always made. Under the assumptions, full knowledge such as, a set of covariates or an underlying causal graph, is typically required. A practical challenge is that in many applications, no such full knowledge or only some partial knowledge is available. In recent years, research has emerged to use search strategies based on graphical causal modelling to discover useful knowledge from data for causal effect estimation, with some mild assumptions, and has shown promise in tackling the practical challenge. In this survey, we review these data-driven methods on causal effect estimation for a single treatment with a single outcome of interest and focus on the challenges faced by data-driven causal effect estimation. We concisely summarise the basic concepts and theories that are essential for data-driven causal effect estimation using graphical causal modelling but are scattered around the literature. We identify and discuss the challenges faced by data-driven causal effect estimation and characterise the existing methods by their assumptions and the approaches to tackling the challenges. We analyse the strengths and limitations of the different types of methods and present an empirical evaluation to support the discussions. We hope this review will motivate more researchers to design better data-driven methods based on graphical causal modelling for the challenging problem of causal effect estimation.
Motivation & Objective
- To address the challenge of unbiased causal effect estimation from observational data when full causal knowledge (e.g., complete causal graph or all confounders) is unavailable.
- To synthesize scattered theoretical foundations of graphical causal modeling relevant to data-driven causal effect estimation, particularly under assumptions like faithfulness and Markov condition.
- To identify and analyze key challenges in data-driven causal effect estimation, including uncertainty in causal structure learning, computational complexity, and latent confounding.
- To evaluate existing methods based on their assumptions, methodological approaches, and performance, with a focus on average treatment effect (ATE) estimation.
- To provide a structured overview and empirical evaluation to guide researchers in selecting and developing better data-driven causal inference methods.
Proposed method
- Systematically reviews graphical causal models (e.g., DAGs, MAGs) as the foundation for representing causal relationships in observational data.
- Categorizes data-driven methods based on how they handle uncertainty in causal structure learning, such as using constraint-based, score-based, or hybrid search strategies.
- Analyzes methods that incorporate instrumental variables (IVs) and conditional IVs to address latent confounding, including approaches like sisVIVE that tolerate invalid instruments.
- Evaluates methods using synthetic and semi-synthetic datasets (e.g., IHDP, Twins) with known ground truth causal effects and structures to assess accuracy and robustness.
- Assesses time complexity and scalability of methods, especially in high-dimensional settings with latent variables.
- Compares methods through empirical evaluation using real-world datasets (e.g., Job training, 401(k), Schoolingreturns) where ground truth is approximated via domain knowledge.
Experimental results
Research questions
- RQ1How do data-driven methods based on graphical causal modeling estimate average treatment effects (ATE) when the full causal structure is unknown or partially observed?
- RQ2What are the key assumptions (e.g., faithfulness, Markov, causal sufficiency) that underlie these methods, and how do they affect estimation accuracy?
- RQ3How do methods handle latent confounders and invalid instrumental variables in causal effect estimation from observational data?
- RQ4What are the trade-offs between methodological complexity, computational efficiency, and estimation accuracy in data-driven causal inference?
- RQ5How can empirical evaluation be meaningfully conducted in the absence of ground-truth causal effects in real-world datasets?
Key findings
- Data-driven causal effect estimation methods based on graphical causal modeling can effectively estimate ATE even with partial knowledge of the causal structure, especially when combined with instrumental variable techniques.
- Methods such as sisVIVE show robustness to the presence of invalid instruments, making them suitable for real-world applications with unmeasured confounders.
- Empirical evaluation on semi-synthetic datasets like IHDP and Twins shows that many data-driven methods achieve reasonable accuracy, though performance varies significantly with data quality and underlying assumptions.
- The lack of ground truth in real-world datasets limits reliable evaluation, and existing benchmarks rely on domain knowledge or empirical estimates as proxies.
- Time complexity remains a significant challenge, particularly for structure learning in high-dimensional settings, limiting scalability of some methods.
- The survey identifies that current methods are still constrained by assumptions like faithfulness and causal sufficiency, and further research is needed to relax these while maintaining estimation validity.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.