Skip to main content
QUICK REVIEW

[Paper Review] Artificial Intelligence for Central Dogma-Centric Multi-Omics: Challenges and Breakthroughs

Lei Xin, Caiyun Huang|arXiv (Cornell University)|Dec 17, 2024
Bioinformatics and Genomic NetworksBiochemistry, Genetics and Molecular Biology3 citations
TL;DR

This paper reviews artificial intelligence (AI) applications in central dogma-centric multi-omics integration, focusing on challenges like high-dimensional noisy data and batch effects, and highlights breakthroughs in foundation models, attention mechanisms, and multi-tissue single-cell data integration. It demonstrates how AI enhances disease subtyping, biomarker discovery, and precision medicine through advanced deep learning and multi-omics data fusion.

ABSTRACT

With the rapid development of high-throughput sequencing platforms, an increasing number of omics technologies, such as genomics, metabolomics, and transcriptomics, are being applied to disease genetics research. However, biological data often exhibit high dimensionality and significant noise, making it challenging to effectively distinguish disease subtypes using a single-omics approach. To address these challenges and better capture the interactions among DNA, RNA, and proteins described by the central dogma, numerous studies have leveraged artificial intelligence to develop multi-omics models for disease research. These AI-driven models have improved the accuracy of disease prediction and facilitated the identification of genetic loci associated with diseases, thus advancing precision medicine. This paper reviews the mathematical definitions of multi-omics, strategies for integrating multi-omics data, applications of artificial intelligence and deep learning in multi-omics, the establishment of foundational models, and breakthroughs in multi-omics technologies, drawing insights from over 130 related articles. It aims to provide practical guidance for computational biologists to better understand and effectively utilize AI-based multi-omics machine learning algorithms in the context of central dogma.

Motivation & Objective

  • To address the limitations of single-omics approaches in linking genotype to phenotype and identifying disease subtypes.
  • To review AI and deep learning techniques for integrating multi-omics data across the central dogma (DNA → RNA → protein).
  • To identify bottlenecks in multi-omics data integration, including batch effects, data sparsity, and computational complexity.
  • To highlight recent advances in foundation models, attention mechanisms, and scalable inference for large biological sequences.
  • To advocate for open, multi-organ multi-omics datasets to accelerate AI-driven discovery in systems biology and precision medicine.

Proposed method

  • Systematic review of 130+ studies on AI in multi-omics, focusing on integration strategies, model architectures, and applications.
  • Application of attention mechanisms and Transformer-based models (e.g., TTT, InstInfer, Mnemosyne) to handle long biological sequences with linear complexity.
  • Use of metric learning and unified embedding techniques (e.g., SCimilarity, CELLama) to enable interpretable, cross-species cell type comparison.
  • Leveraging large-scale single-cell datasets (e.g., 50M+ cells in scFoundation) for pre-training foundation models across tissues and species.
  • Integration of multi-omics data using multi-modal autoencoders and attention-based fusion to model gene regulatory networks.
  • Adoption of 3D parallel inference strategies (e.g., Mnemosyne) to enable real-time processing of up to 10 million sequence contexts.

Experimental results

Research questions

  • RQ1How can AI models effectively integrate multi-omics data across the central dogma to improve disease subtyping and biomarker discovery?
  • RQ2What are the key computational and biological challenges in scaling AI models to long biological sequences and multi-tissue single-cell data?
  • RQ3How do foundation models like scFoundation and CELLama enable cross-species and cross-tissue cell type identification and functional inference?
  • RQ4What role do attention mechanisms and efficient inference architectures play in reducing memory and computational costs for large-scale omics modeling?
  • RQ5How can improved data integration and standardized immune repertoire databases enhance the prediction of immunotherapy response and antigen specificity?

Key findings

  • AI-driven multi-omics integration significantly improves disease prediction accuracy and enables the identification of disease-associated genetic loci and regulatory mechanisms.
  • Foundation models pre-trained on over 50 million single-cell samples (e.g., scFoundation) enable robust cross-tissue and cross-species cell type embedding and functional inference.
  • Models like TTT and Mnemosyne extend sequence modeling to over 16,000 tokens with linear complexity, enabling scalable analysis of long genomic and transcriptomic sequences.
  • Metric learning-based methods such as SCimilarity allow rapid, interpretable queries for cells with similar morphological and transcriptional profiles.
  • Despite progress, only ~10% of untargeted metabolomics data can be structurally annotated, highlighting a major bottleneck in data completeness.
  • The lack of large, publicly available multi-organ multi-omics datasets remains a key barrier to training generalizable AI models, especially in immunology and rare diseases.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.