Skip to main content
QUICK REVIEW

[论文解读] On Learning Causal Structures from Non-Experimental Data without Any Faithfulness Assumption

Hanti Lin, Jiji Zhang|arXiv (Cornell University)|Feb 20, 2018
Bayesian Modeling and Causal Inference参考文献 8被引用 4
一句话总结

本文证明,任何实现最强收敛性的因果学习算法——具体而言,几乎处处收敛结合附着局部一致性的收敛——都必须必然收敛于所有满足忠实性条件的因果贝叶斯网络。若不假设忠实性,这种收敛在普遍意义上是不可能实现的,因此标准实践中针对忠实结构的设计并非可选项,而是实现最优学习性能的逻辑必要条件。

ABSTRACT

Consider the problem of learning, from non-experimental data, the causal (Markov equivalence) structure of the true, unknown causal Bayesian network (CBN) on a given, fixed set of (categorical) variables. This learning problem is known to be so hard that there is no learning algorithm that converges to the truth for all possible CBNs (on the given set of variables). So the convergence property has to be sacrificed for some CBNs---but for which? In response, the standard practice has been to design and employ learning algorithms that secure the convergence property for at least all the CBNs that satisfy the famous faithfulness condition, which implies sacrificing the convergence property for some CBNs that violate the faithfulness condition (Spirtes et al. 2000). This standard design practice can be justified by assuming---that is, accepting on faith---that the true, unknown CBN satisfies the faithfulness condition. But the real question is this: Is it possible to explain, without assuming the faithfulness condition or any of its weaker variants, why it is mandatory rather than optional to follow the standard design practice? This paper aims to answer the above question in the affirmative. We first define an array of modes of convergence to the truth as desiderata that might or might not be achieved by a causal learning algorithm. Those modes of convergence concern (i) how pervasive the domain of convergence is on the space of all possible CBNs and (ii) how uniformly the convergence happens. Then we prove a result to the following effect: for any learning algorithm that tackles the causal learning problem in question, if it achieves the best achievable mode of convergence (considered in this paper), then it must follow the standard design practice of converging to the truth for at least all CBNs that satisfy the faithfulness condition---it is a requirement, not an option.

研究动机与目标

  • 确定在不假设忠实性的情况下,设计因果学习算法以收敛于所有忠实因果贝叶斯网络的标准实践是否合理。
  • 探究是否存在替代学习策略,可在避免牺牲对忠实网络收敛的同时,仍实现强收敛模式。
  • 在非实验数据因果结构学习的背景下,形式化并分析不同的收敛模式(例如,几乎处处收敛、局部一致收敛)。
  • 证明实现最佳可能收敛行为本质上要求对所有忠实因果结构实现收敛,无论是否假设忠实性。
  • 使用密度和拓扑可忽略性等概念,为为何在最优学习性能下无法避免牺牲对忠实网络的收敛提供拓扑学上的解释。

提出的方法

  • 引入一个因果学习的正式框架,其中因果状态表示为一对 (G, P),G 为因果图,P 为联合概率分布。
  • 定义多种收敛模式:几乎处处收敛、局部一致收敛以及附着局部一致收敛,以评估学习算法的性能。
  • 使用拓扑概念——特别是因果状态度量空间中稠密集和开球——分析收敛域的结构。
  • 采用反证法:假设一个学习算法实现了最优收敛,但在某个忠实网络上失败,随后构造一个邻近的因果状态,导致收敛性被违反,从而产生矛盾。
  • 应用关键引理(6.3 和 6.5)表明,稠密且开的收敛域意味着在每个非极小因果状态中都必须牺牲收敛性。
  • 通过使用拓扑上的可忽略性概念(例如,第一纲集)而非勒贝格测度,将结果推广至类别变量之外,使方法适用于连续或无限范围变量。

实验结果

研究问题

  • RQ1是否可以不假设忠实性,而合理化因果发现中针对忠实因果结构的标准实践?
  • RQ2因果学习算法所能实现的最强收敛模式是什么?这些模式对收敛域施加了何种约束?
  • RQ3能否在不收敛于所有忠实因果贝叶斯网络的前提下,实现最优收敛?
  • RQ4因果状态空间的拓扑结构如何影响因果发现中普遍收敛的可行性?
  • RQ5考虑到统计不可识别性的限制,收敛于忠实网络与收敛于非忠实网络之间存在何种必要权衡?

主要发现

  • 任何实现最佳可能收敛模式的因果学习算法——具体而言,几乎处处收敛结合附着局部一致收敛——都必须收敛于所有满足忠实性条件的因果贝叶斯网络。
  • 此类最优算法的收敛域必须在因果状态空间中为稠密且开集,这意味着在每个非极小因果状态中都无法实现收敛。
  • 若此类算法在某个忠实网络上失败,则会产生矛盾,因为这意味着其在该网络的邻域内也必然失败,从而违反了所假设的最优性。
  • 必须收敛于所有忠实网络的要求,并不依赖于是否假设忠实性;它源于因果状态空间的拓扑结构以及所期望的收敛性质。
  • 通过使用拓扑可忽略性概念(如第一纲集)而非测度论概念,结果可推广至连续或无限范围变量。
  • 标准设计实践中针对忠实结构的处理并非出于便利或假设,而是实现因果结构学习中最佳收敛性能的逻辑必然要求。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。