[论文解读] UnPaSt: unsupervised patient stratification by biclustering of omics data
UnPaSt 是一种新颖的无监督双聚类算法,用于多组学数据中的患者分层,旨在通过识别差异表达的双聚类来检测非互斥的疾病亚型。该方法在识别乳腺癌和哮喘中的已知亚型方面优于现有方法,在整体转录组学、单细胞转录组学、蛋白质组学和空间组学数据中揭示了具有生物意义的模式。
Unsupervised patient stratification is essential for disease subtype discovery, yet, despite growing evidence of molecular heterogeneity of non-oncological diseases, popular methods are benchmarked primarily using cancers with mutually exclusive molecular subtypes well-differentiated by numerous biomarkers. Evaluating 22 unsupervised methods, including clustering and biclustering, using simulated and real transcriptomics data revealed their inefficiency in scenarios with non-mutually exclusive subtypes or subtypes discriminated only by few biomarkers. To address these limitations and advance precision medicine, we developed UnPaSt, a novel biclustering algorithm for unsupervised patient stratification based on differentially expressed biclusters. UnPaSt outperformed widely used patient stratification approaches in the de novo identification of known subtypes of breast cancer and asthma. In addition, it detected many biologically insightful patterns across bulk transcriptomics, proteomics, single-cell, spatial transcriptomics, and multi-omics datasets, enabling a more nuanced and interpretable view of high-throughput data heterogeneity than traditionally used methods.
研究动机与目标
- 解决现有无监督方法在非肿瘤性疾病中识别非互斥疾病亚型的局限性。
- 克服在少数特征性生物标志物或亚型重叠场景下,标准聚类和双聚类方法性能不佳的问题。
- 开发一种方法,实现从多样化组学数据类型中可解释的、从头发现生物相关患者亚型。
- 通过实现对复杂疾病中分子异质性的更细致表征,推动精准医学的发展。
提出的方法
- 提出一种双聚类框架,用于在组学数据中识别基因-患者子矩阵(双聚类)中的差异表达。
- 使用统计检验检测在患者亚组中显著富集于差异表达的双聚类。
- 采用贪心优化策略,迭代提取非重叠双聚类,同时保持生物学相关性。
- 通过联合分析转录组学、蛋白质组学和空间转录组学数据,整合多组学层次,以提升亚型分辨率。
- 在多个双聚类集合上采用一致性聚类方法,推导出稳定的患者分组。
- 通过置换检验和错误发现率校正验证双聚类的显著性,以控制假阳性结果。
实验结果
研究问题
- RQ1双聚类方法是否能在非肿瘤性疾病中优于标准聚类方法,以识别非互斥疾病亚型?
- RQ2UnPaSt 在使用无监督学习检测乳腺癌和哮喘等复杂疾病已知亚型方面效果如何?
- RQ3UnPaSt 在包括单细胞和空间转录组学在内的多样化组学数据中,能够多大程度上揭示具有生物意义的模式?
- RQ4在传统方法失效的低生物标志物区分度场景下,UnPaSt 的表现如何?
- RQ5与最先进无监督方法相比,UnPaSt 是否能揭示更具可解释性和生物一致性的患者亚型?
主要发现
- 在乳腺癌和哮喘数据集的已知亚型从头识别中,UnPaSt 超过了22种现有无监督方法。
- 该方法在整体转录组学、蛋白质组学、单细胞转录组学、空间转录组学以及多组学数据中成功检测到具有生物意义的双聚类。
- 即使在差异表达生物标志物较少的场景下,UnPaSt 仍能识别出有意义的患者亚组,而标准方法则失败。
- 该算法在多种组学模态中揭示了可解释的分子模式,增强了疾病异质性的分辨率。
- 在基准评估中,与基于聚类的方法相比,UnPaSt 在检测非互斥亚型方面表现出更优的鲁棒性和敏感性。
- UnPaSt 识别出的双聚类在下游富集分析中显著富集于已知疾病通路,表现出强烈的功能一致性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。