[Paper Review] On Learning Causal Structures from Non-Experimental Data without Any Faithfulness Assumption
This paper establishes that any causal learning algorithm achieving the strongest possible convergence to the true causal structure—specifically, almost everywhere convergence combined with adherent local uniformity—must necessarily converge for all causal Bayesian networks satisfying the faithfulness condition. Without assuming faithfulness, such convergence is impossible to achieve universally, making the standard practice of targeting faithful structures not optional but logically required for optimal learning performance.
Consider the problem of learning, from non-experimental data, the causal (Markov equivalence) structure of the true, unknown causal Bayesian network (CBN) on a given, fixed set of (categorical) variables. This learning problem is known to be so hard that there is no learning algorithm that converges to the truth for all possible CBNs (on the given set of variables). So the convergence property has to be sacrificed for some CBNs---but for which? In response, the standard practice has been to design and employ learning algorithms that secure the convergence property for at least all the CBNs that satisfy the famous faithfulness condition, which implies sacrificing the convergence property for some CBNs that violate the faithfulness condition (Spirtes et al. 2000). This standard design practice can be justified by assuming---that is, accepting on faith---that the true, unknown CBN satisfies the faithfulness condition. But the real question is this: Is it possible to explain, without assuming the faithfulness condition or any of its weaker variants, why it is mandatory rather than optional to follow the standard design practice? This paper aims to answer the above question in the affirmative. We first define an array of modes of convergence to the truth as desiderata that might or might not be achieved by a causal learning algorithm. Those modes of convergence concern (i) how pervasive the domain of convergence is on the space of all possible CBNs and (ii) how uniformly the convergence happens. Then we prove a result to the following effect: for any learning algorithm that tackles the causal learning problem in question, if it achieves the best achievable mode of convergence (considered in this paper), then it must follow the standard design practice of converging to the truth for at least all CBNs that satisfy the faithfulness condition---it is a requirement, not an option.
Motivation & Objective
- To determine whether the standard practice of designing causal learning algorithms to converge on all faithful causal Bayesian networks is justifiable without assuming faithfulness.
- To investigate whether there exist alternative learning strategies that avoid sacrificing convergence on faithful networks while still achieving strong convergence modes.
- To formalize and analyze different modes of convergence (e.g., almost everywhere, locally uniform) in the context of causal structure learning from non-experimental data.
- To prove that achieving the best possible convergence behavior inherently requires convergence on all faithful causal structures, regardless of whether faithfulness is assumed.
- To provide a topological justification for why sacrificing convergence on faithful networks is inescapable for optimal learning performance, using concepts like density and topological negligibility.
Proposed method
- Introduces a formal framework for causal learning where causal states are represented as pairs (G, P), with G being a causal graph and P a joint probability distribution.
- Defines multiple modes of convergence: almost everywhere convergence, locally uniform convergence, and adherent local uniformity, to evaluate learning algorithm performance.
- Uses topological concepts—specifically, dense sets and open balls in a metric space of causal states—to analyze the structure of the domain of convergence.
- Employs a proof by contradiction: assumes a learning algorithm achieves optimal convergence but fails on some faithful network, then constructs a nearby causal state that violates convergence, leading to contradiction.
- Applies key lemmas (6.3 and 6.5) to show that dense and open domains of convergence imply that convergence must be sacrificed in every non-minimal causal state.
- Generalizes results beyond categorical variables by using topological notions of negligibility (e.g., meager sets) instead of Lebesgue measure, enabling applicability to continuous or infinite-range variables.
Experimental results
Research questions
- RQ1Is it possible to justify the standard practice of targeting faithful causal structures in causal discovery without assuming faithfulness?
- RQ2What are the strongest possible convergence modes that a causal learning algorithm can achieve, and what constraints do they impose on the domain of convergence?
- RQ3Can a learning algorithm achieve optimal convergence without converging on all faithful causal Bayesian networks?
- RQ4How does the topological structure of the space of causal states affect the feasibility of universal convergence in causal discovery?
- RQ5What is the necessary trade-off between convergence on faithful networks and convergence on unfaithful ones, given the limitations imposed by statistical nonidentifiability?
Key findings
- Any causal learning algorithm that achieves the best possible convergence mode—specifically, almost everywhere convergence combined with adherent local uniformity—must converge on all causal Bayesian networks that satisfy the faithfulness condition.
- The domain of convergence for such an optimal algorithm must be dense and open in the space of causal states, which implies that convergence cannot be achieved in every non-minimal causal state.
- A contradiction arises if such an algorithm fails to converge on a faithful network, as it would imply failure to converge in a neighborhood of that network, violating the assumed optimality.
- The necessity of converging on all faithful networks is not contingent on assuming faithfulness; it follows from the topological structure of the space of causal states and the desired convergence properties.
- The results generalize beyond categorical variables to continuous and infinite-range variables by using topological notions of negligibility (e.g., meager sets) instead of measure-theoretic ones.
- The standard design practice of targeting faithful structures is not a matter of convenience or assumption, but a logical requirement for achieving optimal convergence performance in causal structure learning.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.