[Paper Review] Perturbations and Causality in Gaussian Latent Variable Models
This paper proposes DirectLikelihood, a maximum-likelihood estimator for identifying causal structures in Gaussian latent variable models using unspecific perturbation data, without requiring the response or latent variables to remain unperturbed. It leverages system-wide invariances to uniquely recover the population causal structure, extending to non-Gaussian linear models.
Causal inference is a challenging problem with observational data alone. The task becomes easier when having access to data from perturbing the underlying system, even when happening in a non-randomized way: this is the setting we consider, encompassing also latent confounding variables. To identify causal relations among a collections of covariates and a response variable, existing procedures rely on at least one of the following assumptions: i) the response variable remains unperturbed, ii) the latent variables remain unperturbed, and iii) the latent effects are dense. In this paper, we examine a perturbation model for interventional data, which can be viewed as a mixed-effects linear structural causal model, over a collection of Gaussian variables that does not satisfy any of these conditions. We propose a maximum-likelihood estimator -- dubbed DirectLikelihood -- that exploits system-wide invariances to uniquely identify the population causal structure from unspecific perturbation data, and our results carry over to linear structural causal models without requiring Gaussianity. We illustrate the utility of our framework on synthetic data as well as real data involving California reservoirs and protein expressions.
Motivation & Objective
- To address causal inference in observational data with interventional perturbations, especially when standard assumptions do not hold.
- To overcome limitations of existing methods that require the response or latent variables to remain unperturbed.
- To develop a method that identifies causal structures without assuming dense latent effects or intervention-specific design.
- To establish a framework valid for both Gaussian and non-Gaussian linear structural causal models using system-wide invariances.
- To provide a robust estimator—DirectLikelihood—that uniquely identifies the population causal structure from unspecific perturbation data.
Proposed method
- Proposes a mixed-effects linear structural causal model as a perturbation framework for Gaussian variables with latent confounders.
- Develops a maximum-likelihood estimator, DirectLikelihood, that exploits invariances across multiple perturbed interventions.
- Uses system-wide invariance properties to identify causal effects even when both observed and latent variables are perturbed.
- Derives the likelihood function under the mixed-effects model and optimizes it to estimate causal coefficients.
- Extends the method to non-Gaussian linear structural causal models by leveraging the same invariance principles.
- Validates the estimator through theoretical analysis and empirical evaluation on synthetic and real-world data.
Experimental results
Research questions
- RQ1Can causal structure be uniquely identified from unspecific perturbation data when neither the response nor latent variables are assumed to remain unperturbed?
- RQ2How can system-wide invariances in perturbed data be exploited to estimate causal effects without assuming dense latent effects?
- RQ3Does the proposed DirectLikelihood estimator maintain identifiability and consistency under general perturbation regimes in Gaussian latent variable models?
- RQ4To what extent does the method generalize to non-Gaussian linear structural causal models?
- RQ5How does the performance of DirectLikelihood compare to existing methods in settings violating standard assumptions?
Key findings
- DirectLikelihood uniquely identifies the population causal structure from unspecific perturbation data without requiring the response or latent variables to be unperturbed.
- The method achieves identifiability even when latent effects are sparse, overcoming a key limitation of prior approaches.
- The estimator is consistent and asymptotically efficient under the assumed mixed-effects model framework.
- Empirical results on synthetic data confirm the method’s ability to recover true causal structures under various perturbation regimes.
- Applications to California reservoir data and protein expression data demonstrate the method’s practical utility and robustness in real-world settings.
- The framework generalizes beyond Gaussianity, maintaining validity for linear structural causal models with non-Gaussian errors.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.