[Paper Review] Disentangling regional impacts of joint teleconnections using causal representation learning
This paper introduces DAG-VAE, a causal representation learning method that embeds a physics-informed DAG in the latent space of a variational autoencoder to jointly learn nonlinear reduced representations of teleconnections and their causal effects on Greater Horn of Africa rainfall, and to generate data-driven counterfactuals.
Understanding teleconnections of large-scale modes of climate variability is relevant for seasonal predictability and support a dynamical understanding of climatic changes. While numerical model experiments are the most common approach for investigating counterfactual climate responses, their conclusions are subject to model biases. Data-driven approaches offer a complementary perspective. Deep learning can extract reduced-dimensional patterns but usually lacks causal interpretability, while causal methods can disentangle signals in the presence of confounding yet are typically based on simple indices. Treating dimensionality reduction and causal inference separately thereby risks losing the teleconnection signal of interest. This paper introduces DAG-VAE, a causal representation learning approach that embeds a physics-informed directed acyclic graph in the latent space of a variational autoencoder. Combining deep learning with causal inference, the method jointly learns nonlinear reduced representations of large-scale modes of variability and their causal interactions. We apply DAG-VAE to disentangle the influences of the Pacific and Indian Oceans on the short rains over the Greater Horn of Africa. Trained on seasonal hindcasts, the method identifies dynamically meaningful representations and recovers spatial response patterns consistent with SST-replacement experiments. Trained on reanalysis data, DAG-VAE identifies a different response pattern to direct influence of the tropical Pacific, highlighting potential model biases and the value of DAG-VAE as a complementary, data-driven approach for estimating spatial causal response patterns from observations. Finally, we demonstrate the ability of the method to generate data-driven counterfactuals of extreme short rain seasons, with potential applications for forecast-based early action and scenario planning.
Motivation & Objective
- Develop a data-driven method to jointly learn low-dimensional representations of large-scale climate variability and their causal interactions.
- Estimate spatial causal response patterns of precipitation to individual teleconnection drivers.
- Enable generation of data-driven counterfactual scenarios for extreme short-rain seasons to support forecast-based action and planning.
Proposed method
- Embed a physics-informed directed acyclic graph (DAG) in the latent space of a variational autoencoder (VAE) to model causal relations among latent representations of tropical Pacific SSTs, Indian Ocean SSTs, and Greater Horn of Africa precipitation.
- Use a sparsity-regularized, linear latent-space causal model within a nonlinear VAE framework to ensure identifiability and disentanglement of causal factors.
- Train DAG-VAE on SEAS5 seasonal hindcasts and ERA5 reanalysis to learn interpretable latent representations and their causal connections.
- Interrogate the model with latent-space interventions to estimate direct causal effects and to generate counterfactual precipitation patterns.
- Compare against PCA-based and index-based baselines to evaluate reconstruction, predictive skill, and robustness of learned causal factors.
- Benchmark against SST-replacement experiments from prior work to assess consistency of causal pathways and spatial response patterns.
Experimental results
Research questions
- RQ1Can a DAG-informed latent space within a VAE disentangle the joint causal influences of Pacific and Indian Ocean SST anomalies on GHA short rains?
- RQ2Do the learned latent representations yield spatially interpretable precipitation response patterns consistent with causal theory and SST-replacement experiments?
- RQ3Can the method generate plausible data-driven counterfactuals of extreme short-rain seasons based on identified causal drivers?
Key findings
- DAG-VAE identifies nonlinear, physically meaningful latent representations for tropical Pacific SSTs, Indian Ocean SSTs, and GHA precipitation that reflect ENSO and IOD patterns.
- The method achieves higher anomaly correlation (r ≈ 0.68) and total precipitation R^2 (≈ 0.58) for GHA rainfall than PCA and index-based baselines.
- Interventions in the latent space recover known SST replacement results: positive IOD directly causes wet anomalies across the GHA; ENSO directly causes an east–west dipole with inland drying and offshore wetting.
- DAG-VAE shows robust latent representations with high mean correlation coefficients across multiple trainings, supporting identifiability in practice.
- On reanalysis data, the direct Pacific influence yields different spatial patterns than on hindcasts, highlighting model biases and the value of a data-driven approach.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.