[Paper Review] AIRIVA: A Deep Generative Model of Adaptive Immune Repertoires
AIRIVA is a deep generative model that learns disentangled, interpretable latent representations of T-cell receptor (TCR) repertoires to disentangle systematic biases like disease status, HLA type, and sequencing depth. It enables robust prediction of disease associations, generation of in-silico counterfactual repertoires via latent intervention, and identification of disease-specific TCRs validated with external data, demonstrating improved performance on COVID-19 and HSV-1/2 repertoires.
Recent advances in immunomics have shown that T-cell receptor (TCR) signatures can accurately predict active or recent infection by leveraging the high specificity of TCR binding to disease antigens. However, the extreme diversity of the adaptive immune repertoire presents challenges in reliably identifying disease-specific TCRs. Population genetics and sequencing depth can also have strong systematic effects on repertoires, which requires careful consideration when developing diagnostic models. We present an Adaptive Immune Repertoire-Invariant Variational Autoencoder (AIRIVA), a generative model that learns a low-dimensional, interpretable, and compositional representation of TCR repertoires to disentangle such systematic effects in repertoires. We apply AIRIVA to two infectious disease case-studies: COVID-19 (natural infection and vaccination) and the Herpes Simplex Virus (HSV-1 and HSV-2), and empirically show that we can disentangle the individual disease signals. We further demonstrate AIRIVA's capability to: learn from unlabelled samples; generate in-silico TCR repertoires by intervening on the latent factors; and identify disease-associated TCRs validated using TCR annotations from external assay data.
Motivation & Objective
- To address the challenge of identifying disease-specific TCRs in highly diverse, sparse, and systematically biased adaptive immune repertoires.
- To develop a semi-supervised generative model that learns disentangled, interpretable latent factors linked to biological variables such as disease labels, HLA type, and sequencing depth.
- To enable in-silico generation of counterfactual TCR repertoires by intervening on latent factors for model validation and TCR-disease association discovery.
- To improve model robustness and generalization by leveraging unlabelled repertoires, especially in low-sample-size, heterogeneous cohorts common in immunomics.
- To validate the model’s ability to identify biologically meaningful TCR-disease associations using external TCR annotation data.
Proposed method
- AIRIVA is a variational autoencoder with a factorized conditional prior that enforces disentanglement of latent factors corresponding to biological variables like disease status and HLA type.
- It uses a classifier-guided loss to encourage disentanglement by aligning latent factors with predictive labels while maintaining generative capacity.
- The model is trained on TCR count data with partial labels, enabling semi-supervised learning from unlabelled repertoires to improve generalization.
- Counterfactual repertoires are generated by intervening on specific latent factors (e.g., toggling disease status) in the latent space and decoding to synthetic TCR repertoires.
- Latent-space projections and counterfactual consistency metrics are used to validate disentanglement and guide model selection.
- The framework supports conditional generation and CATE (Conditional Average Treatment Effect) estimation to identify TCRs associated with specific disease states.
Experimental results
Research questions
- RQ1Can a deep generative model learn disentangled, interpretable representations of TCR repertoires that separate biological factors like disease status from technical confounders such as sequencing depth?
- RQ2To what extent can AIRIVA improve disease prediction performance by leveraging unlabelled TCR repertoires in low-data regimes?
- RQ3Can AIRIVA generate biologically plausible in-silico counterfactual repertoires by manipulating latent factors, and are these consistent with known TCR-disease associations?
- RQ4How well can AIRIVA identify disease-specific TCRs, and can these predictions be validated using external TCR annotation data?
- RQ5Does disentangled representation learning in AIRIVA enhance model robustness across diverse subgroups, such as distinguishing natural infection from vaccination in COVID-19?
Key findings
- AIRIVA successfully disentangles disease signals in TCR repertoires, demonstrating improved robustness in distinguishing natural SARS-CoV-2 infection from vaccination status in COVID-19.
- The model achieves improved label prediction performance by leveraging unlabelled repertoires, highlighting its utility in low-sample-size, heterogeneous immunomics cohorts.
- Generated counterfactual repertoires are consistent across interventions, validating the model’s ability to simulate realistic biological scenarios through latent space manipulation.
- AIRIVA identifies TCRs associated with spike and non-spike antigens in SARS-CoV-2, with predictions confirmed using external TCR annotation data.
- The disentangled latent factors enable interpretable diagnosis, longitudinal monitoring, and identification of TCRs for cellular therapy development.
- AIRIVA demonstrates scalability potential for multi-label settings, including infectious disease panels, HLA typing, and batch effect detection, with a strong foundation in disentanglement and counterfactual generation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.