[Paper Review] A Survey of Learning Causality with Data: Problems and Methods
This survey provides a structured review of traditional and frontier methods for learning causality from data, comparing SCMs and the potential outcomes framework, and detailing data types, problems, and methods in causal inference and discovery, with connections to machine learning in the big-data era.
This work considers the question of how convenient access to copious data impacts our ability to learn causal effects and relations. In what ways is learning causality in the era of big data different from -- or the same as -- the traditional one? To answer this question, this survey provides a comprehensive and structured review of both traditional and frontier methods in learning causality and relations along with the connections between causality and machine learning. This work points out on a case-by-case basis how big data facilitates, complicates, or motivates each approach.
Motivation & Objective
- Clarify how big data changes the learning of causal effects and causal relations compared to traditional settings.
- Systematically categorize data types and corresponding causal problems (effects vs. relations).
- Present core frameworks (structural causal models and potential outcomes) and their use for identification and estimation.
- Survey traditional and frontier methods for causal inference and discovery under observational data.
- Discuss connections between causality learning and machine learning disciplines such as supervised learning, domain adaptation, and reinforcement learning.
Proposed method
- Introduce two foundational causal frameworks: structural causal models (SCMs) and the potential outcome framework (POF).
- Explain key concepts: causal graphs, d-separation, back-door criterion, do-calculus, and interventional vs. observational distributions.
- Classify data for causal learning into three types: data for causal effects (with/without confounders, big data considerations), data for causal relations, and time-series settings.
- Outline standard identification strategies to eliminate confounding bias, including adjustment via back-door criteria and alternative identification when back-door fails.
- Describe how SCMs and POF are used to define and estimate average treatment effects (ATE), individual treatment effects (ITE), and conditional effects (CATE).
- Discuss limitations of do-calculus (e.g., individual-level queries and counterfactuals) and introduce counterfactual notation in Pearl’s framework.
- Summarize how ground truth for causal effects/relations is obtained (randomized experiments, simulations, prior knowledge).
- Provide guidance on how data types map to problems and methods for learning causal effects and causal relations.
Experimental results
Research questions
- RQ1How does the era of big data affect our ability to learn causal effects and causal relations?
- RQ2What data types and problem formulations enable learning causal effects and causal relations from observational data?
- RQ3What are the foundational frameworks and identification strategies used to estimate causal quantities from data?
- RQ4How do connections between causality and machine learning inform methodological developments?
- RQ5What open problems remain in learning causality with big data?
Key findings
- Big data can mitigate some confounding via richer feature sets but introduces new biases and high-dimensional challenges.
- Causal questions split into learning causal effects (inference) and learning causal relations (discovery), each requiring different data and methods.
- SCMs and the potential outcomes framework are logically equivalent and provide complementary views for identification and estimation.
- Identification often relies on back-door adjustment or other strategies when back-door criteria hold; when they do not, alternative identification methods are discussed.
- Ground truth for causal effects is typically obtained from randomized experiments or domain knowledge simulations, with counterfactuals posing additional challenges.
- The survey connects causal learning to machine learning topics like supervised/semi-supervised learning, domain adaptation, and reinforcement learning.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.