Skip to main content
QUICK REVIEW

[Paper Review] A technical review of multi-omics data integration methods: from classical statistical to deep generative approaches

Ana R. Baião, Zhaoxiang Cai|arXiv (Cornell University)|Jan 29, 2025
Bioinformatics and Genomic Networks5 citations
TL;DR

A comprehensive review of multi-omics data integration methods, with emphasis on deep generative models (especially variational autoencoders) for imputation, joint embedding, and batch correction, plus discussions of loss functions, regularisation, and future directions.

ABSTRACT

The rapid advancement of high-throughput sequencing and other assay technologies has resulted in the generation of large and complex multi-omics datasets, offering unprecedented opportunities for advancing precision medicine strategies. However, multi-omics data integration presents significant challenges due to the high dimensionality, heterogeneity, experimental gaps, and frequency of missing values across data types. Computational methods have been developed to address these issues, employing statistical and machine learning approaches to uncover complex biological patterns and provide deeper insights into our understanding of disease mechanisms. Here, we comprehensively review state-of-the-art multi-omics data integration methods with a focus on deep generative models, particularly variational autoencoders (VAEs) that have been widely used for data imputation and augmentation, joint embedding creation, and batch effect correction. We explore the technical aspects of loss functions and regularisation techniques including adversarial training, disentanglement and contrastive learning. Moreover, we discuss recent advancements in foundation models and the integration of emerging data modalities, while describing the current limitations and outlining future directions for enhancing multi-modal methodologies in biomedical research.

Motivation & Objective

  • Motivate the challenges of integrating high-dimensional, heterogeneous multi-omics data for precision medicine.
  • Survey state-of-the-art methods with emphasis on deep generative models for imputation, augmentation, and joint embedding.
  • Discuss technical aspects such as loss functions, regularisation techniques, and training strategies.
  • Highlight recent advances in foundation models and new data modalities for biomedical integration.
  • Identify current limitations and propose future research directions in multi-modal biomedical data integration.

Proposed method

  • Survey of classical statistical and machine learning approaches for multi-omics integration.
  • Emphasis on deep generative models, especially variational autoencoders (VAEs), for imputation, augmentation, and joint embedding.
  • Discussion of loss functions and regularisation techniques including adversarial training, disentanglement, and contrastive learning.
  • Analysis of training strategies and model architectures used for cross-modal data integration.
  • Consideration of foundation models and emerging data modalities in multi-omics.
  • Critical appraisal of limitations and future directions.

Experimental results

Research questions

  • RQ1What are the key methodological advancements in multi-omics data integration across classical statistics, ML, and deep generative approaches?
  • RQ2How have VAEs and related generative models been applied to imputation, augmentation, and joint embedding in multi-omics data?
  • RQ3What loss functions, regularisation strategies, and training paradigms best address high dimensionality and heterogeneity in multi-omics data?
  • RQ4What are the limitations of current methods, and what future directions are promising for integrating new data modalities in biomedicine?
  • RQ5How do foundation models and emerging modalities impact multi-omics data integration?

Key findings

  • Deep generative models, notably VAEs, are widely used for data imputation, augmentation, and joint embedding in multi-omics integration.
  • Adversarial training, disentanglement, and contrastive learning are important regularisation techniques discussed in the context of multi-omics VAEs.
  • Recent advances include foundations models and integration of emerging data modalities.
  • The review synthesizes technical aspects of loss functions and regularisation, highlighting their role in addressing high dimensionality and data heterogeneity.
  • Limitations of current approaches and gaps point to future directions for more robust, scalable, and interpretable multi-modal methods.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.