[Paper Review] Conditionally Invariant Representation Learning for Disentangling Cellular Heterogeneity
The paper introduces a conditionally invariant deep generative model that disentangles invariant biological signals from domain-specific noise to improve multi-domain single-cell data integration and interpretation.
This paper presents a novel approach that leverages domain variability to learn representations that are conditionally invariant to unwanted variability or distractors. Our approach identifies both spurious and invariant latent features necessary for achieving accurate reconstruction by placing distinct conditional priors on latent features. The invariant signals are disentangled from noise by enforcing independence which facilitates the construction of an interpretable model with a causal semantic. By exploiting the interplay between data domains and labels, our method simultaneously identifies invariant features and builds invariant predictors. We apply our method to grand biological challenges, such as data integration in single-cell genomics with the aim of capturing biological variations across datasets with many samples, obtained from different conditions or multiple laboratories. Our approach allows for the incorporation of specific biological mechanisms, including gene programs, disease states, or treatment conditions into the data integration process, bridging the gap between the theoretical assumptions and real biological applications. Specifically, the proposed approach helps to disentangle biological signals from data biases that are unrelated to the target task or the causal explanation of interest. Through extensive benchmarking using large-scale human hematopoiesis and human lung cancer data, we validate the superiority of our approach over existing methods and demonstrate that it can empower deeper insights into cellular heterogeneity and the identification of disease cell states.
Motivation & Objective
- Motivate learning representations that separate invariant biological signals from domain-specific noise across multi-domain single-cell datasets.
- Propose a conditionally invariant generative model that identifies both spurious and invariant latent factors.
- Provide identifiability guarantees and validate on large-scale hematopoiesis and lung cancer scRNA-seq data.
Proposed method
- Use a variational autoencoder with a conditionally factorized prior to achieve identifiability of latent variables.
- Split latent space into invariant (Z_I) and spurious (Z_S) components to capture stable vs. domain-varying information.
- Impose independence between Z_I and Z_S to enable reconstruction while isolating invariant features.
- Incorporate auxiliary sample information (d) and environment (e) to model dependencies and guide disentanglement.
- Aim for an invariant predictor that uses Z_I to predict labels Y with performance across environments.
- Compare with NF-iVAE and other invariant learning approaches to benchmark identifiability and data integration.

Experimental results
Research questions
- RQ1How can latent representations be partitioned into invariant and non-invariant components in multi-domain single-cell data?
- RQ2Can a conditionally identifiable VAE separate biological signals from technical or domain-driven noise while preserving predictive performance across environments?
- RQ3What priors and assumptions enable identifiability and practical applicability to single-cell genomics across datasets?
Key findings
- The proposed method identifies both invariant and spurious latent variables in a conditional VAE framework.
- The model is identifiable up to simple transformations and a permutation of latent variables.
- Evaluations on large-scale human hematopoiesis and human lung cancer scRNA-seq data with 49 samples across two cancer types show improved data integration and cell-state profiling (as per the text).
- The approach enables incorporation of biological mechanisms such as gene programs, disease states, or treatment conditions into the data integration process.
- The method demonstrates superiority over existing invariant and identifiable deep generative models for single-cell data integration and cell-type annotation.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.