Skip to main content
QUICK REVIEW

[Paper Review] Identifiability Guarantees for Causal Disentanglement from Soft Interventions

Jiaqi Zhang, Chandler Squires|arXiv (Cornell University)|Jul 12, 2023
Gene expression and cancer classificationBiochemistry, Genetics and Molecular Biology9 citations
TL;DR

The paper proves identifiability of latent causal variables and structure from unpaired observational and soft-interventional data under a generalized faithfulness assumption, even when the latent variables are unobserved, and presents an AVB-based learning algorithm.

ABSTRACT

Causal disentanglement aims to uncover a representation of data using latent variables that are interrelated through a causal model. Such a representation is identifiable if the latent model that explains the data is unique. In this paper, we focus on the scenario where unpaired observational and interventional data are available, with each intervention changing the mechanism of a latent variable. When the causal variables are fully observed, statistically consistent algorithms have been developed to identify the causal model under faithfulness assumptions. We here show that identifiability can still be achieved with unobserved causal variables, given a generalized notion of faithfulness. Our results guarantee that we can recover the latent causal model up to an equivalence class and predict the effect of unseen combinations of interventions, in the limit of infinite data. We implement our causal disentanglement framework by developing an autoencoding variational Bayes algorithm and apply it to the problem of predicting combinatorial perturbation effects in genomics.

Motivation & Objective

  • Motivate causal disentanglement as learning a latent causal representation that supports interventions.
  • Show identifiability of latent causal models when latent variables are unobserved under a generalized faithfulness notion.
  • Provide theoretical results identifying ancestral relations and the full causal structure up to a CD-equivalence class.
  • Develop a practical learning algorithm to recover the CD-equivalence class from data.
  • Demonstrate applicability to genomics by predicting effects of unseen genetic perturbations.

Proposed method

  • Model X as f(U) with U following a DAG G and unobserved latent U; interventions modify P(U_i | pa_G(i)) to form P^I(U_i | pa_G(i)).
  • Assume polynomial, full-rank mixing function f and support conditions to identify U up to linear transformations (Assumption 1).
  • Establish identifiability of G and interventions up to the CD-equivalence class using generalized faithfulness (Assumptions 1–3) and interventional data (Theorems 1–2).
  • Use transitive closure TS(G) to identify ancestral relations and then refine to the CD-equivalence class of (G, I1,…,IK).
  • Propose a discrepancy-based variational autoencoder (VAE) with a deep structural causal model decoder to learn U, G, and I from data; enable counterfactual/interventional sampling.
  • Extend framework to predict combinatorial unseen interventions via learned U and G.

Experimental results

Research questions

  • RQ1Can latent causal structure and intervention targets be identified from unpaired observational and soft-interventional data when latent variables are unobserved?
  • RQ2Under what conditions (Assumptions 1–3) can the causal graph and interventions be identified up to CD-equivalence?
  • RQ3How can one recover ancestral relations and full direct edges from interventional data with linear mixing?
  • RQ4Can we learn a scalable algorithm to estimate the CD-equivalence class from data and predict unseen intervention effects?
  • RQ5How does the approach apply to high-dimensional biological data, such as genomics perturbations?

Key findings

  • Identifiability is achievable with unobserved latent variables under a generalized faithfulness notion (Theorems 1–2).
  • The latent causal model is identifiable up to a CD-equivalence class, enabling prediction of unseen intervention combinations.
  • Interventional data allows identification of ancestral relations and, under Assumption 3, direct edges in many cases.
  • A gradient-based AVB approach (DiscrepancyVAE) can learn the CD-equivalence class via a deep SCM decoder and sparsity-promoting objectives.
  • The framework is demonstrated on genomics data for predicting combinatorial perturbation effects.
  • The method supports extrapolation to unseen combinations of interventions through the learned latent structure.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.