[Paper Review] Causal Machine Learning: A Survey and Open Problems
This paper surveys Causal Machine Learning (CausalML), classifies work into five problem areas, reviews applications and benchmarks, and outlines open problems across causality, learning, explanations, fairness, and reinforcement learning.
Causal Machine Learning (CausalML) is an umbrella term for machine learning methods that formalize the data-generation process as a structural causal model (SCM). This perspective enables us to reason about the effects of changes to this process (interventions) and what would have happened in hindsight (counterfactuals). We categorize work in CausalML into five groups according to the problems they address: (1) causal supervised learning, (2) causal generative modeling, (3) causal explanations, (4) causal fairness, and (5) causal reinforcement learning. We systematically compare the methods in each category and point out open problems. Further, we review data-modality-specific applications in computer vision, natural language processing, and graph representation learning. Finally, we provide an overview of causal benchmarks and a critical discussion of the state of this nascent field, including recommendations for future work.
Motivation & Objective
- Provide a minimal, self-contained introduction to causality for ML researchers.
- Taxonomize CausalML work into five problem classes and compare approaches within each class.
- Review modality-specific applications (vision, NLP, graph learning) and causal benchmarks.
- Discuss practical benefits, challenges, and open research questions in CausalML.
- Offer recommendations for future work and field directions.
Proposed method
- Introduce key causality concepts and standard formalisms (Bayesian networks, graphical causal models, SCMs).
- Define interventions (do-operator) and counterfactuals, and explain the ladder of causation.
- Describe Independent Mechanisms and structural assumptions underlying SCMs.
- Provide a taxonomy of CausalML categories: causal supervised learning, causal generative modeling, causal explanations, causal fairness, and causal reinforcement learning.
- Survey applications in computer vision, natural language processing, and graph representation learning.
- Summarize causal benchmarks and the state of the field, including good/bad/ugly considerations.
Experimental results
Research questions
- RQ1What are the core causality formalisms and how do they operationalize interventions and counterfactuals?
- RQ2How can we categorize existing CausalML methods and compare approaches within each category?
- RQ3What are the key applications and benchmarks driving CausalML across modalities?
- RQ4What open problems and future directions emerge across causality, learning, explanations, fairness, and RL?
Key findings
- The paper provides a minimal, self-contained introduction to causality for ML researchers.
- It proposes a taxonomy dividing CausalML work into five problem classes and systematically compares approaches within each class.
- It reviews modality-specific applications (vision, NLP, graph learning) and causal benchmarks.
- It discusses the pros (good) and challenges (bad/ugly) of adopting CausalML in practice.
- It highlights open problems and offers guidance for future research directions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.