[Paper Review] UnPaSt: unsupervised patient stratification by biclustering of omics data
UnPaSt is a novel unsupervised biclustering algorithm for patient stratification in multi-omics data, designed to detect non-mutually exclusive disease subtypes by identifying differentially expressed biclusters. It outperforms existing methods in identifying known subtypes in breast cancer and asthma, revealing biologically meaningful patterns across bulk and single-cell transcriptomics, proteomics, and spatial omics data.
Unsupervised patient stratification is essential for disease subtype discovery, yet, despite growing evidence of molecular heterogeneity of non-oncological diseases, popular methods are benchmarked primarily using cancers with mutually exclusive molecular subtypes well-differentiated by numerous biomarkers. Evaluating 22 unsupervised methods, including clustering and biclustering, using simulated and real transcriptomics data revealed their inefficiency in scenarios with non-mutually exclusive subtypes or subtypes discriminated only by few biomarkers. To address these limitations and advance precision medicine, we developed UnPaSt, a novel biclustering algorithm for unsupervised patient stratification based on differentially expressed biclusters. UnPaSt outperformed widely used patient stratification approaches in the de novo identification of known subtypes of breast cancer and asthma. In addition, it detected many biologically insightful patterns across bulk transcriptomics, proteomics, single-cell, spatial transcriptomics, and multi-omics datasets, enabling a more nuanced and interpretable view of high-throughput data heterogeneity than traditionally used methods.
Motivation & Objective
- Address the limitation of existing unsupervised methods in identifying non-mutually exclusive disease subtypes in non-oncological diseases.
- Overcome the poor performance of standard clustering and biclustering methods in scenarios with few discriminative biomarkers or overlapping subtypes.
- Develop a method that enables interpretable, de novo discovery of biologically relevant patient subtypes from diverse omics data types.
- Improve precision medicine by enabling more nuanced characterization of molecular heterogeneity in complex diseases.
Proposed method
- Proposes a biclustering framework that identifies differentially expressed gene-by-patient submatrices (biclusters) in omics data.
- Uses statistical testing to detect biclusters significantly enriched for differential expression across patient subgroups.
- Applies a greedy refinement strategy to iteratively extract non-overlapping biclusters while preserving biological relevance.
- Integrates multiple omics layers by jointly analyzing transcriptomic, proteomic, and spatial transcriptomic datasets to enhance subtype resolution.
- Employs a consensus clustering approach over multiple bicluster sets to derive stable patient groupings.
- Validates bicluster significance using permutation testing and false discovery rate correction to control for false positives.
Experimental results
Research questions
- RQ1Can biclustering methods outperform standard clustering in identifying non-mutually exclusive disease subtypes in non-oncological diseases?
- RQ2How effective is UnPaSt in detecting known subtypes of complex diseases such as breast cancer and asthma using unsupervised learning?
- RQ3To what extent can UnPaSt uncover biologically meaningful patterns in diverse omics data, including single-cell and spatial transcriptomics?
- RQ4How does UnPaSt perform in low-biomarker-discrimination scenarios where traditional methods fail?
- RQ5Can UnPaSt reveal more interpretable and biologically coherent patient subtypes compared to state-of-the-art unsupervised methods?
Key findings
- UnPaSt outperformed 22 existing unsupervised methods in de novo identification of known subtypes in breast cancer and asthma datasets.
- The method successfully detected biologically relevant biclusters in bulk transcriptomics, proteomics, single-cell, spatial transcriptomics, and multi-omics data.
- UnPaSt identified meaningful patient subgroups even in scenarios with few differentially expressed biomarkers, where standard methods failed.
- The algorithm revealed interpretable molecular patterns across diverse omics modalities, enhancing the resolution of disease heterogeneity.
- In benchmark evaluations, UnPaSt demonstrated superior robustness and sensitivity in detecting non-mutually exclusive subtypes compared to clustering-based approaches.
- The biclusters identified by UnPaSt were enriched for known disease pathways and showed strong functional coherence in downstream enrichment analysis.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.