Skip to main content
QUICK REVIEW

[论文解读] Bayesian Inference for Tumor Subclones Accounting for Sequencing and Structural Variants

Ju‐Hee Lee, Peter Müeller|arXiv (Cornell University)|Sep 25, 2014
Cancer Genomics and Diagnostics参考文献 27被引用 5
一句话总结

该论文提出了一种贝叶斯特征分配模型,通过下一代测序数据联合推断肿瘤亚克隆拷贝数、体细胞变异等位基因计数和细胞分数。通过使用基于非参数先验的三个随机矩阵(L、Z、w)对亚克隆结构进行建模,该方法在模拟和真实肺癌数据上表现出色,成功检测到两个推断出的亚克隆,显著提升了亚克隆结构推断的准确性。

ABSTRACT

Tumor samples are heterogeneous. They consist of different subclones that are characterized by differences in DNA nucleotide sequences and copy numbers on multiple loci. Heterogeneity can be measured through the identification of the subclonal copy number and sequence at a selected set of loci. Understanding that the accurate identification of variant allele fractions greatly depends on a precise determination of copy numbers, we develop a Bayesian feature allocation model for jointly calling subclonal copy numbers and the corresponding allele sequences for the same loci. The proposed method utilizes three random matrices, L, Z and w to represent subclonal copy numbers (L), numbers of subclonal variant alleles (Z) and cellular fractions of subclones in samples (w), respectively. The unknown number of subclones implies a random number of columns for these matrices. We use next-generation sequencing data to estimate the subclonal structures through inference on these three matrices. Using simulation studies and a real data analysis, we demonstrate how posterior inference on the subclonal structure is enhanced with the joint modeling of both structure and sequencing variants on subclonal genomes. Software is available at http://compgenome.org/BayClone2.

研究动机与目标

  • 为解决缺乏联合估计肿瘤样本中亚克隆拷贝数和体细胞变异等位基因计数的计算模型的问题。
  • 通过在统一框架中整合结构变异(拷贝数)和序列变异(SNV),提升体细胞变异等位基因分数估计的准确性。
  • 提供一种概率推理方法,用于亚克隆结构推断,以考虑拷贝数和等位基因计数中的不确定性。
  • 通过揭示异质性肿瘤的真实亚克隆组成,实现更精确的癌症预后和靶向治疗。
  • 开发一种灵活且可扩展的模型,可进一步整合SNP芯片数据或患者水平聚类等额外数据源。

提出的方法

  • 该方法使用三个潜在随机矩阵:L用于亚克隆拷贝数,Z用于体细胞变异等位基因计数,w用于亚克隆细胞分数。
  • 采用基于有限版本的多项式印度餐厅帘过程的非参数先验,以建模未知的亚克隆数量,允许矩阵具有随机的列维度。
  • 采用马尔可夫链蒙特卡洛(MCMC)采样对三个矩阵进行后验推断,利用测序读数计数和观察到的体细胞变异频率。
  • 通过整合各位点的测序深度和等位基因特异性信号,联合估计拷贝数状态和体细胞变异等位基因计数。
  • 通过模拟研究调整超参数,并利用合成数据和真实肺癌NGS数据验证后验估计结果。
  • 该框架支持不确定性量化,并可扩展以整合额外数据类型,如SNP芯片数据。

实验结果

研究问题

  • RQ1联合建模拷贝数变异和单核苷酸变异在多大程度上能提升亚克隆结构推断的准确性?
  • RQ2在异质性肿瘤样本中,整合拷贝数信息对体细胞变异等位基因分数估计有何影响?
  • RQ3贝叶斯非参数模型能否在不预先假设固定亚克隆数量的情况下,有效推断亚克隆数量及其遗传特征?
  • RQ4与PyClone等现有工具相比,该方法在估计亚克隆流行率和聚类位点方面表现如何?
  • RQ5该模型在多大程度上提升了下游应用,如个性化癌症治疗和临床试验设计?

主要发现

  • 在真实肺癌数据集中,该模型成功推断出两个亚克隆(C*=2),与有限空间异质性的生物预期一致。
  • 后验预测检查显示,估计的读数计数(N̂st)和变异分数(p̂st)集中在观测值附近,表明模型对数据的拟合良好。
  • 估计的拷贝数矩阵L*在许多位点显示为三个拷贝,反映了数据中观察到的高测序深度。
  • 细胞分数矩阵w*在四个空间上接近的肿瘤样本间表现出高度相似性,支持有限亚克隆多样性在生物学上的合理性。
  • 与PyClone相比,该方法在联合推断亚克隆结构方面表现更优,因为PyClone仅基于经验体细胞变异分数聚类位点,而未直接建模亚克隆群体。
  • 在模拟研究中,该模型表现出稳健性,在不同噪声水平和亚克隆复杂度下均能准确恢复亚克隆结构。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。