Skip to main content
QUICK REVIEW

[论文解读] Identifying cancer subtypes in glioblastoma by combining genomic, transcriptomic and epigenomic data

Richard S. Savage, Zoubin Ghahramani|arXiv (Cornell University)|Apr 12, 2013
Gene expression and cancer classification参考文献 10被引用 7
一句话总结

本研究提出了一种增强的非参数贝叶斯方法,用于整合基因组、转录组和表观基因组数据,以识别胶质母细胞瘤亚型。通过联合分析来自277例TCGA胶质母细胞瘤样本的基因表达、拷贝数变异、甲基化和微RNA数据,该方法识别出8个共识亚型,其中甲基化数据强烈预测复发(log-rank p = 2.0×10⁻³),包括一个具有显著低甲基化水平的亚型,其10年内复发率接近零。

ABSTRACT

We present a nonparametric Bayesian method for disease subtype discovery in multi-dimensional cancer data. Our method can simultaneously analyse a wide range of data types, allowing for both agreement and disagreement between their underlying clustering structure. It includes feature selection and infers the most likely number of disease subtypes, given the data. We apply the method to 277 glioblastoma samples from The Cancer Genome Atlas, for which there are gene expression, copy number variation, methylation and microRNA data. We identify 8 distinct consensus subtypes and study their prognostic value for death, new tumour events, progression and recurrence. The consensus subtypes are prognostic of tumour recurrence (log-rank p-value of $3.6 imes 10^{-4}$ after correction for multiple hypothesis tests). This is driven principally by the methylation data (log-rank p-value of $2.0 imes 10^{-3}$) but the effect is strengthened by the other 3 data types, demonstrating the value of integrating multiple data types. Of particular note is a subtype of 47 patients characterised by very low levels of methylation. This subtype has very low rates of tumour recurrence and no new events in 10 years of follow up. We also identify a small gene expression subtype of 6 patients that shows particularly poor survival outcomes. Additionally, we note a consensus subtype that showly a highly distinctive data signature and suggest that it is therefore a biologically distinct subtype of glioblastoma. The code is available from https://sites.google.com/site/multipledatafusion/

研究动机与目标

  • 开发一种稳健的方法,用于从多维分子数据中识别癌症亚型,尤其在不同类型数据可能不一致的情况下。
  • 在不事先假设亚型数量的前提下推断最优的疾病亚型数目。
  • 在多个数据类型(基因表达、拷贝数、甲基化、微RNA)中同时执行特征选择。
  • 评估通过跨数据类型共识聚类推导出的亚型的预后价值。
  • 通过分析不同类型数据间聚类结构的不一致性,探索胶质母细胞瘤的生物学复杂性。

提出的方法

  • 采用基于狄利克雷过程混合模型的非参数贝叶斯框架,实现无需预先指定亚型数量的灵活聚类。
  • 在MDI算法基础上扩展为数据特定模型:对连续数据(如基因表达、拷贝数)使用高斯分布,对二值数据(如甲基化、微RNA)使用多项分布。
  • 通过估计特征包含的后验概率,实现每类数据的特征选择,仅保留最具有信息量的基因/探针。
  • 采用结合吉布斯抽样与分裂-合并操作的混合MCMC采样器,以提高高维聚类中的混合效率与收敛性。
  • 通过后验相似性矩阵聚合不同数据类型中的聚类分配结果进行共识聚类,对聚类标签进行排序以揭示结构。
  • 通过phi_kl参数的后验均值量化不同类型数据间的共识程度,该参数用于估计不同类型间共享的聚类结构。

实验结果

研究问题

  • RQ1统一的统计模型能否通过整合基因组、转录组和表观基因组数据,识别出具有生物意义的胶质母细胞瘤亚型?
  • RQ2不同分子数据类型在胶质母细胞瘤样本聚类中的一致性或不一致性程度如何?
  • RQ3哪种分子数据类型对所识别亚型的预后能力贡献最大?
  • RQ4是否存在具有独特生物学特征的亚型,例如极端甲基化水平或不良生存结局?
  • RQ5共识聚类方法是否能揭示比单一数据类型更具有生物学相关性的亚型?

主要发现

  • 共识聚类识别出8个不同的胶质母细胞瘤亚型,其中甲基化数据是肿瘤复发最强的预测因子(log-rank p = 2.0×10⁻³)。
  • 47名患者组成的低甲基化亚型在10年随访期内未出现肿瘤复发或新发事件。
  • 6名患者的基因表达亚型表现出极差的生存结局,提示存在一种罕见且侵袭性强的胶质母细胞瘤形式。
  • 8名患者的共识亚型具有高度独特的多组学特征,复发率低,可能代表一种生物学上独立的胶质母细胞瘤亚型。
  • 四种数据类型之间的聚类结构仅部分重叠,表明其潜在生物学机制比单一统一的亚型结构更为复杂。
  • 该方法成功在无需先验假设的情况下推断出亚型数量并选择相关特征,展示了在多组学整合中具备稳健性与生物学相关性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。