Skip to main content
QUICK REVIEW

[论文解读] Causal Discovery from Heterogeneous/Nonstationary Data with Independent Changes

Biwei Huang, Kun Zhang|arXiv (Cornell University)|Mar 5, 2019
Bayesian Modeling and Causal Inference参考文献 45被引用 14
一句话总结

本文提出CD-NOD,一种基于异质性或非平稳数据中跨领域或时间的独立变化的非参数因果发现框架。它通过条件独立性检验识别因果骨架,利用分布漂移推断因果方向,并估计机制变化背后的低维‘驱动因素’——即使在存在混淆和非平稳性的情况下也能实现稳健的因果推断。

ABSTRACT

It is commonplace to encounter heterogeneous or nonstationary data, of which the underlying generating process changes across domains or over time. Such a distribution shift feature presents both challenges and opportunities for causal discovery. In this paper, we develop a framework for causal discovery from such data, called Constraint-based causal Discovery from heterogeneous/NOnstationary Data (CD-NOD), to find causal skeleton and directions and estimate the properties of mechanism changes. First, we propose an enhanced constraint-based procedure to detect variables whose local mechanisms change and recover the skeleton of the causal structure over observed variables. Second, we present a method to determine causal orientations by making use of independent changes in the data distribution implied by the underlying causal model, benefiting from information carried by changing distributions. After learning the causal structure, next, we investigate how to efficiently estimate the "driving force" of the nonstationarity of a causal mechanism. That is, we aim to extract from data a low-dimensional representation of changes. The proposed methods are nonparametric, with no hard restrictions on data distributions and causal mechanisms, and do not rely on window segmentation. Furthermore, we find that data heterogeneity benefits causal structure identification even with particular types of confounders. Finally, we show the connection between heterogeneity/nonstationarity and soft intervention in causal discovery. Experimental results on various synthetic and real-world data sets (task-fMRI and stock market data) are presented to demonstrate the efficacy of the proposed methods.

研究动机与目标

  • 解决由于跨领域或随时间变化的机制导致的分布漂移所带来的因果发现挑战。
  • 开发一种不依赖窗口分割或严格分布假设的非参数方法。
  • 识别机制发生变化的变量,并从未观测到的非平稳数据中恢复因果骨架。
  • 利用数据分布中的独立变化来定向因果边,从而提升因果方向识别的准确性。
  • 通过再生核希尔伯特空间(RKHS)中变化表示的低秩逼近,估计代表机制变化根本原因的低维‘驱动因素’。

提出的方法

  • 提出一种基于核的条件独立性检验的约束性方法,用于检测机制发生变化的变量并恢复因果骨架。
  • 采用核分布嵌入(KDE)表示联合分布,并以非参数方式测试条件独立性。
  • 引入独立变化原则,通过在变量间独立的分布漂移来定向因果边。
  • 使用去混淆集合Z来评估边缘分布与条件分布之间的条件独立性,以推断因果方向。
  • 通过再生核希尔伯特空间(RKHS)中变化表示的低秩逼近,估计非平稳性的‘驱动因素’。
  • 将框架扩展至处理同时包含瞬时和滞后因果效应的动态系统,并考虑平稳混淆因子的影响。

实验结果

研究问题

  • RQ1如何从机制在不同领域或时间点发生变化的异质性或非平稳数据中可靠地恢复因果骨架?
  • RQ2能否利用机制独立变化引起的分布漂移来识别因果方向,而无需干预?
  • RQ3如何估计机制变化根本原因的低维、可解释的表示(即‘驱动因素’)?
  • RQ4即使存在未观测到的混淆因子,数据异质性是否仍能提升因果结构识别能力?
  • RQ5非平稳性与因果发现中的软干预之间有何关联?

主要发现

  • 所提出的CD-NOD框架仅使用观测数据,成功从异质性和非平稳数据中恢复了因果骨架及其方向。
  • 该方法通过利用分布的独立变化来识别因果方向,在合成数据和真实世界数据(包括fMRI和股票市场数据)中均实现了高精度。
  • 驱动因素估计方法提取了机制变化的低维表示,能够捕捉非平稳性的根本原因。
  • 该框架在各种混淆场景下保持稳健,包括平稳的未观测混淆因子,表明即使存在混淆,数据异质性仍可辅助因果发现。
  • 理论分析证实,在满足三个条件之一(撞衫、单一变化或涉及边的变化)时,基于独立变化原则可识别因果方向。
  • 在任务fMRI和股票市场数据上的实证结果表明,CD-NOD在因果结构恢复和机制变化检测方面优于基线方法。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。