Skip to main content
QUICK REVIEW

[论文解读] Identifiability Assumptions and Algorithm for Directed Graphical Models with Feedback

Gunwoong Park, Garvesh Raskutti|arXiv (Cornell University)|Feb 14, 2016
Bayesian Modeling and Causal Inference参考文献 22被引用 4
一句话总结

本文为有向循环图(DCG)模型引入了两种新的可识别性假设——最大 d-分离规则(MDR)和最稀疏马尔可夫表示(SMR),作为对忠实性假设的更弱替代。MDR 假设在理论上被证明严格弱于忠实性,且在模拟中展现出更可靠的骨架恢复性能,优于基于忠实性的算法在小规模 DCG 模型上的表现。

ABSTRACT

Directed graphical models provide a useful framework for modeling causal or directional relationships for multivariate data. Prior work has largely focused on identifiability and search algorithms for directed acyclic graphical (DAG) models. In many applications, feedback naturally arises and directed graphical models that permit cycles occur. In this paper we address the issue of identifiability for general directed cyclic graphical (DCG) models satisfying the Markov assumption. In particular, in addition to the faithfulness assumption which has already been introduced for cyclic models, we introduce two new identifiability assumptions, one based on selecting the model with the fewest edges and the other based on selecting the DCG model that entails the maximum number of d-separation rules. We provide theoretical results comparing these assumptions which show that: (1) selecting models with the largest number of d-separation rules is strictly weaker than the faithfulness assumption; (2) unlike for DAG models, selecting models with the fewest edges does not necessarily result in a milder assumption than the faithfulness assumption. We also provide connections between our two new principles and minimality assumptions. We use our identifiability assumptions to develop search algorithms for small-scale DCG models. Our simulation study supports our theoretical results, showing that the algorithms based on our two new principles generally out-perform algorithms based on the faithfulness assumption in terms of selecting the true skeleton for DCG models.

研究动机与目标

  • 解决有向循环图(DCG)模型中的模型可识别性挑战,这类模型常见于具有反馈回路的系统中。
  • 克服忠实性假设的局限性,该假设过于严格,在实践中(尤其是数据有限时)很少成立。
  • 提出更温和但理论基础坚实的可识别性假设,以确保 DCG 模型的恢复。
  • 基于新假设设计搜索算法,并与最先进方法进行性能对比。
  • 通过模拟证明,所提出的算法比基于忠实性的方法更可靠地恢复真实图骨架。

提出的方法

  • 将 DAG 中的最稀疏马尔可夫表示(SMR)和节俭性假设适配至 DCG 模型,提出作为新的可识别性标准。
  • 引入最大 d-分离规则(MDR)假设,即选择能蕴含最多 d-分离规则的 DCG 模型。
  • 理论分析对比了 MDR、SMR、忠实性和最小性假设,表明 MDR 严格弱于忠实性但强于 P-最小性。
  • 开发了基于 MDR 和 SMR 假设的搜索算法(算法 1),通过穷举搜索恢复 DCG 的马尔可夫等价类。
  • 使用 Fisher 的条件相关性检验(α = 0.001)估计数据中的条件独立性,作为骨架恢复的基础。
  • 在小规模 DCG 模型上评估性能,使用不同样本量(n ∈ {100, 200, 500, 1000})和期望邻域大小(1 至 4)的合成数据。

实验结果

研究问题

  • RQ1在 DCG 模型中,最大 d-分离规则(MDR)假设是否严格弱于忠实性假设?
  • RQ2SMR 和 MDR 假设在强度上与忠实性和最小性假设相比如何?
  • RQ3基于 MDR 和 SMR 假设的算法是否能比基于忠实性的算法更可靠地恢复 DCG 的真实骨架?
  • RQ4MDR 和 SMR 基础算法在小规模 DCG 模型上的实际性能,与 FCI+ 和 GES 等最先进方法相比如何?
  • RQ5满足每种可识别性假设(CFC、MDR、SMR、P-最小性)的 DCG 模型比例,如何随样本大小和图密度变化?

主要发现

  • MDR 假设严格弱于忠实性假设,意味着其适用于更广泛的 DCG 模型类别。
  • SMR 假设在识别模型的集合上,强于 P-最小性但弱于忠实性。
  • 模拟结果表明,基于 MDR 和 SMR 的算法在所有样本量和邻域大小下,骨架恢复准确率均优于 FCI+ 算法。
  • 对于密集图,GES 算法表现更优,因其偏好密集结构,但其无法正确恢复循环图。
  • 满足 MDR 假设的模型比例大于满足忠实性假设的模型比例,但小于满足 P-最小性假设的模型比例,证实了假设强度的理论排序。
  • 随着样本量增加,满足每种假设的模型比例上升,这是由于条件独立性检验的误差减少所致。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。