Skip to main content
QUICK REVIEW

[论文解读] Topological Pressure and Coding Sequence Density Estimation in the Human Genome

David Koslicki, Daniel J. Thompson|arXiv (Cornell University)|Sep 27, 2011
RNA and protein synthesis mechanisms参考文献 37被引用 1
一句话总结

本论文提出了一种基于拓扑压强(topological pressure)的新方法,用于估算脊椎动物和无脊椎动物基因组中的编码序列(CDS)密度。该方法基于人类基因组数据训练了64组三联体权重,在66,000 bp窗口范围内对小鼠(Mus musculus)、猕猴(Rhesus macaque)和黑腹果蝇(Drosophila melanogaster)的基因组进行CDS密度预测,结果合理,并利用学习到的权重成功区分了人类基因组中750–5000 bp序列的外显子与内含子。

ABSTRACT

We give a new approach to coding sequence (CDS) density estimation in genomic analysis based on the topological pressure, which we develop from a well known concept in ergodic theory. Topological pressure measures the weighted information content of a finite word, and incorporates 64 parameters which can be interpreted as a choice of weight for each nucleotide triplet. We train the parameters so that the topological pressure fits the observed coding sequence density on the human genome, and use this to give ab initio predictions of CDS density over windows of size around 66,000bp on the genomes of Mus Musculus, Rhesus Macaque and Drososphilia Melanogaster. While the differences between these genomes are too great to expect that training on the human genome could predict, for example, the exact locations of genes, we demonstrate that our method gives reasonable estimates for the coarse scale problem of predicting CDS density. Inspired again by ergodic theory, the weightings of the nucleotide triplets obtained from our training procedure are used to define a probability distribution on finite sequences, which can be used to distinguish between intron and exon sequences from the human genome of lengths between 750bp and 5,000bp. At the end of the paper, we explain the theoretical underpinning for our approach, which is the theory of Thermodynamic Formalism from the dynamical systems literature. Mathematica and MATLAB implementations of our method are available at this http URL.

研究动机与目标

  • 开发一种新的从头预测方法,用于在多种物种的大基因组窗口中估算编码序列(CDS)密度。
  • 应用遍历理论和热力学形式化理论中的概念,建立基因组序列统计的模型。
  • 利用拓扑压强训练三联体特异性权重,以拟合人类基因组中观察到的CDS密度。
  • 将训练好的模型推广至非人类基因组,尽管存在进化分歧。
  • 利用学习到的权重分布,对人类基因组中750–5000 bp的基因组区域进行外显子与内含子的分类。

提出的方法

  • 将拓扑压强定义为有限基因组词上的加权信息度量,其中64个参数对应所有可能的核苷酸三联体。
  • 利用人类基因组中观察到的CDS密度数据,通过最小化66,000 bp窗口内的预测误差,对64组三联体权重进行训练。
  • 将训练好的拓扑压强模型应用于预测小鼠(Mus musculus)、猕猴(Rhesus macaque)和黑腹果蝇(Drosophila melanogaster)基因组中的CDS密度。
  • 利用学习到的三联体权重定义有限序列上的概率分布,从而实现外显子与内含子区域的统计区分。
  • 借助热力学形式化理论的理论框架,为模型的构建和稳定性提供理论支持。
  • 在Mathematica和MATLAB中实现该方法,以确保可复现性并实现公众访问。

实验结果

研究问题

  • RQ1在人类CDS密度数据上训练的拓扑压强能否有效预测其他远缘关系基因组中的CDS密度?
  • RQ2基于人类数据训练得到的三联体权重模型,在基因组结构不同的非人类物种中,其泛化能力如何?
  • RQ3从拓扑压强学习到的权重分布是否能有效区分人类基因组中750–5000 bp长度的外显子与内含子序列?
  • RQ4使用拓扑压强进行基因组序列分析的理论基础是什么?
  • RQ5该模型在不依赖已知基因注释的情况下,能在多大程度上捕捉粗粒度的编码潜力?

主要发现

  • 在人类基因组上训练的拓扑压强模型,可在小鼠(Mus musculus)、猕猴(Rhesus macaque)和黑腹果蝇(Drosophila melanogaster)的66,000 bp窗口中合理预测CDS密度。
  • 训练得到的三联体权重可构建序列上的概率分布,有效区分人类基因组中长度为750–5,000 bp的外显子与内含子区域。
  • 该方法表明,拓扑压强可作为基因组序列分析中一种稳健的信息论度量,超越简单的序列组成分析。
  • 热力学形式化理论提供的理论基础,为模型的构建和行为提供了严格的数学依据。
  • 该模型实现了无需依赖先前基因注释或比对数据的从头CDS密度预测。
  • 已在Mathematica和MATLAB中公开提供实现代码,支持方法的复现与进一步拓展。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。