Skip to main content
QUICK REVIEW

[Paper Review] Subtype-Former: a deep learning approach for cancer subtype discovery with multi-omics data

Hai Yang, Yuhang Sheng|arXiv (Cornell University)|Jul 28, 2022
Bioinformatics and Genomic Networks4 citations
TL;DR

Subtype-Former proposes a deep learning framework combining MLP and Transformer blocks to extract low-dimensional representations from multi-omics data for improved cancer subtype discovery. Evaluated on TCGA 10 cancer types, it outperforms state-of-the-art methods in survival-based subtyping and identifies 50 key biomarkers for precision oncology.

ABSTRACT

Motivation: Cancer is heterogeneous, affecting the precise approach to personalized treatment. Accurate subtyping can lead to better survival rates for cancer patients. High-throughput technologies provide multiple omics data for cancer subtyping. However, precise cancer subtyping remains challenging due to the large amount and high dimensionality of omics data. Results: This study proposed Subtype-Former, a deep learning method based on MLP and Transformer Block, to extract the low-dimensional representation of the multi-omics data. K-means and Consensus Clustering are also used to achieve accurate subtyping results. We compared Subtype-Former with the other state-of-the-art subtyping methods across the TCGA 10 cancer types. We found that Subtype-Former can perform better on the benchmark datasets of more than 5000 tumors based on the survival analysis. In addition, Subtype-Former also achieved outstanding results in pan-cancer subtyping, which can help analyze the commonalities and differences across various cancer types at the molecular level. Finally, we applied Subtype-Former to the TCGA 10 types of cancers. We identified 50 essential biomarkers, which can be used to study targeted cancer drugs and promote the development of cancer treatments in the era of precision medicine.

Motivation & Objective

  • To address the challenge of accurate cancer subtyping in the face of high-dimensional, heterogeneous multi-omics data.
  • To develop a deep learning model capable of integrating diverse omics data (e.g., genomics, transcriptomics) for robust molecular subtyping.
  • To improve survival prediction accuracy by identifying biologically meaningful cancer subtypes using multi-omics integration.
  • To enable pan-cancer analysis by discovering conserved molecular patterns across different cancer types.
  • To identify a set of 50 essential biomarkers with potential for targeted therapy development in precision medicine.

Proposed method

  • Employs a hybrid architecture combining Multilayer Perceptron (MLP) and Transformer blocks to model complex patterns in multi-omics data.
  • Uses attention mechanisms within the Transformer blocks to capture long-range dependencies and feature interactions across omics layers.
  • Applies dimensionality reduction via learned representations to distill high-dimensional omics data into low-dimensional, informative embeddings.
  • Utilizes K-means and Consensus Clustering on the learned representations to derive stable and biologically interpretable cancer subtypes.
  • Integrates multiple omics data types (e.g., DNA methylation, gene expression, copy number variation) into a unified representation space.
  • Validates subtyping stability and biological relevance through survival analysis and cross-cancer comparison.

Experimental results

Research questions

  • RQ1Can a deep learning model combining MLP and Transformer architectures improve the accuracy of cancer subtyping using multi-omics data?
  • RQ2How does Subtype-Former perform in comparison to state-of-the-art methods on benchmark TCGA datasets with over 5,000 tumors?
  • RQ3To what extent can Subtype-Former identify biologically coherent and prognostically distinct subtypes across multiple cancer types?
  • RQ4What are the key molecular features (biomarkers) that distinguish the identified subtypes, and are they relevant for targeted therapy?
  • RQ5Can the model reveal conserved molecular patterns across different cancers in a pan-cancer analysis?

Key findings

  • Subtype-Former outperformed existing state-of-the-art methods in cancer subtyping accuracy on TCGA datasets with more than 5,000 tumor samples.
  • The model achieved superior survival prediction performance, indicating that the identified subtypes are biologically and clinically meaningful.
  • In pan-cancer analysis, Subtype-Former revealed shared molecular features across different cancer types, highlighting conserved biological pathways.
  • The method identified 50 essential biomarkers across the 10 TCGA cancer types, which may serve as potential targets for precision cancer therapies.
  • Consensus clustering applied to the learned representations produced stable and reproducible subtypes, enhancing confidence in the results.
  • The integration of MLP and Transformer components enabled effective representation learning from high-dimensional, heterogeneous omics data.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.