Skip to main content
QUICK REVIEW

[論文レビュー] Topological Pressure and Coding Sequence Density Estimation in the Human Genome

David Koslicki, Daniel J. Thompson|arXiv (Cornell University)|Sep 27, 2011
RNA and protein synthesis mechanisms参考文献 37被引用数 1
ひとこと要約

この論文は、エルゴディック理論および熱力学的形式主義の概念であるトポロジカル圧力に基づき、脊椎動物および無脊椎動物ゲノムにおけるコード領域(CDS)密度を推定する新しい手法を紹介する。64組のトリプレット重みをヒトゲノムデータで学習させることで、マウス、サル、ドーパミラのゲノムにおいて66,000 bpのウィンドウ単位でCDS密度を合理的な精度で予測し、学習された重みを用いてヒトゲノム配列(750–5000 bp)におけるエクソンとイントロンを区別する。

ABSTRACT

We give a new approach to coding sequence (CDS) density estimation in genomic analysis based on the topological pressure, which we develop from a well known concept in ergodic theory. Topological pressure measures the weighted information content of a finite word, and incorporates 64 parameters which can be interpreted as a choice of weight for each nucleotide triplet. We train the parameters so that the topological pressure fits the observed coding sequence density on the human genome, and use this to give ab initio predictions of CDS density over windows of size around 66,000bp on the genomes of Mus Musculus, Rhesus Macaque and Drososphilia Melanogaster. While the differences between these genomes are too great to expect that training on the human genome could predict, for example, the exact locations of genes, we demonstrate that our method gives reasonable estimates for the coarse scale problem of predicting CDS density. Inspired again by ergodic theory, the weightings of the nucleotide triplets obtained from our training procedure are used to define a probability distribution on finite sequences, which can be used to distinguish between intron and exon sequences from the human genome of lengths between 750bp and 5,000bp. At the end of the paper, we explain the theoretical underpinning for our approach, which is the theory of Thermodynamic Formalism from the dynamical systems literature. Mathematica and MATLAB implementations of our method are available at this http URL.

研究の動機と目的

  • 多様な種において、大規模なゲノムウィンドウにおけるコード領域(CDS)密度を推定するための新しいア・ビ・ニシオ手法を開発すること。
  • エルゴディック理論および熱力学的形式主義の概念を用いて、ゲノム配列統計をモデル化すること。
  • ヒトゲノムにおける観察されたCDS密度に適合させるために、トポロジカル圧力に基づく64組のトリプレット固有重みを学習すること。
  • 進化的に分岐した非ヒトゲノムに対しても、学習済みモデルを一般化してCDS密度を予測できること。
  • 学習された重み分布を用いて、ヒトゲノム配列(750–5000 bp)における遺伝子領域をエクソンとイントロンに分類すること。

提案手法

  • 64組のトリプレットに対応する全核酸トリプレットの重み付き情報測度としてトポロジカル圧力を定義する。
  • ヒトゲノムの観察されたCDS密度データを用いて、66,000 bpウィンドウにおける予測誤差を最小化するように64組のトリプレット重みを学習する。
  • 学習済みのトポロジカル圧力モデルを用いて、マウス(Mus musculus)、サル(Rhesus macaque)、ドーパミラ(Drosophila melanogaster)ゲノムにおけるCDS密度を予測する。
  • 学習されたトリプレット重みを用いて有限シーケンス上での確率分布を定義し、エクソンとイントロン領域の統計的区別を可能にする。
  • 熱力学的形式主義の理論的枠組みを活用して、モデルの構築および安定性の正当化を図る。
  • 再現性および公開アクセスを目的として、MathematicaおよびMATLABの両方で手法を実装する。

実験結果

リサーチクエスチョン

  • RQ1ヒトのCDS密度データで学習したトポロジカル圧力は、系統的に遠く離れた他のゲノムにおけるCDS密度を予測できるか?
  • RQ2ヒトデータから得たトリプレット重みに基づくモデルは、ゲノム構造が異なる非ヒト種に対し、どの程度一般化できるか?
  • RQ3トポロジカル圧力から学習された重み分布は、ヒトゲノム領域におけるエクソンとイントロン配列を効果的に区別できるか?
  • RQ4トポロジカル圧力をゲノム配列解析に用いることの理論的根拠は何か?
  • RQ5このモデルは、既知の遺伝子アノテーションに依存せずに、粗いスケールのコード化可能性をどの程度捉えられるか?

主な発見

  • ヒトゲノムで学習したトポロジカル圧力モデルは、マウス、サル、ドーパミラのゲノムにおいて66,000 bpウィンドウ単位でCDS密度を合理的な精度で予測できた。
  • 学習済みのトリプレット重みは、750 bpから5,000 bpの長さを有するヒトのエクソンおよびイントロン領域を効果的に区別できる確率分布を定義した。
  • この手法は、トポロジカル圧力が単なる塩基組成を超えた情報理論的測度として、ゲノム配列解析に強固に適用可能であることを示した。
  • 熱力学的形式主義に基づく理論的基盤により、モデルの構築および挙動に数学的厳密性が与えられた。
  • このモデルは、事前の遺伝子アノテーションやアラインメントデータを一切必要とせず、ア・ビ・ニシオでCDS密度を予測した。
  • MathematicaおよびMATLABによる公開実装が提供されており、手法の再現および拡張が可能である。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。