Skip to main content
QUICK REVIEW

[Paper Review] AI-driven multi-omics integration for multi-scale predictive modeling of causal genotype-environment-phenotype relationships

You Wu, Lei Xie|arXiv (Cornell University)|Jul 8, 2024
Gene expression and cancer classification8 citations
TL;DR

Proposes an AI-powered, biology-inspired framework to integrate multi-omics data across biological scales and species to predict causal genotype-environment-phenotype relationships under perturbations. It reviews perturbation omics resources and surveys state-of-the-art multi-omics integration methods.

ABSTRACT

Despite the wealth of single-cell multi-omics data, it remains challenging to predict the consequences of novel genetic and chemical perturbations in the human body. It requires knowledge of molecular interactions at all biological levels, encompassing disease models and humans. Current machine learning methods primarily establish statistical correlations between genotypes and phenotypes but struggle to identify physiologically significant causal factors, limiting their predictive power. Key challenges in predictive modeling include scarcity of labeled data, generalization across different domains, and disentangling causation from correlation. In light of recent advances in multi-omics data integration, we propose a new artificial intelligence (AI)-powered biology-inspired multi-scale modeling framework to tackle these issues. This framework will integrate multi-omics data across biological levels, organism hierarchies, and species to predict causal genotype-environment-phenotype relationships under various conditions. AI models inspired by biology may identify novel molecular targets, biomarkers, pharmaceutical agents, and personalized medicines for presently unmet medical needs.

Motivation & Objective

  • Motivate the need to predict phenotypes from genotypes under environmental perturbations using endophenotypes as connecting links.
  • Propose a biology-inspired AI framework that integrates multi-omics data across scales and species to infer causal G-E-P relationships.
  • Survey perturbation omics data resources and current machine learning approaches for multi-omics integration to identify limitations and opportunities.

Proposed method

  • Review perturbation omics data resources (e.g., TCGA, LINCS, DepMap, scPerturb, PharmacoDB, ProteomicsDB) and their applicability to G-E-P modeling.
  • Summarize and critique state-of-the-art unsupervised, supervised, and knowledge-graph-based multi-omics integration methods (autoencoders, transformers, contrastive learning, GNNs, etc.).
  • Highlight biology-inspired AI modeling principles for cross-level, cross-scale, cross-species data integration aimed at predicting phenotypic responses to unprecedented perturbations.
Figure 1: Illustration of cross-level, cross-scale, cross-species multi-omics data integration
Figure 1: Illustration of cross-level, cross-scale, cross-species multi-omics data integration

Experimental results

Research questions

  • RQ1What are the available perturbation omics data resources and how can they support predictive G-E-P modeling?
  • RQ2What are the strengths and limitations of current unsupervised, supervised, and graph-based multi-omics integration methods for cross-level genotype-environment-phenotype prediction?
  • RQ3How can AI models inspired by biology enable cross-species translation of molecular mechanisms to human phenotypes under perturbations?
  • RQ4What gaps exist in data, generalization, and causality that the proposed framework should address?

Key findings

  • There is a wealth of perturbation omics data across modalities and organisms, but labeled data for phenotype prediction under perturbations remains scarce.
  • Current methods span unsupervised, supervised, and knowledge-graph approaches, each with trade-offs related to data pairing requirements, modality alignment, and cross-domain generalizability.
  • Biology-inspired AI frameworks that integrate multi-omics across scales and species hold promise for identifying novel targets, biomarkers, and personalized therapies.
  • Foundational models (e.g., transformers) and cross-species analyses (e.g., GeneCompass) illustrate potential for cross-domain and cross-species G-E-P insights, albeit with limitations such as reliance on paired data and domain alignment challenges.
Figure 2: Illustration of multi-modal supervised learning. (a) A conventional strategy that requires paired data for all the modalities simultaneously. (b) An end-to-end deep neural network explicitly models asymmetric information flows from DNAs to RNAs to proteins to metabolites and ultimately to
Figure 2: Illustration of multi-modal supervised learning. (a) A conventional strategy that requires paired data for all the modalities simultaneously. (b) An end-to-end deep neural network explicitly models asymmetric information flows from DNAs to RNAs to proteins to metabolites and ultimately to

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.