Skip to main content
QUICK REVIEW

[論文レビュー] Multiscale persistent functions for biomolecular structure characterization

Kelin Xia, Zhiming Li|arXiv (Cornell University)|Dec 26, 2016
Topological and Geometric Data Analysis参考文献 54被引用数 4
ひとこと要約

本稿では、多スケール剛性関数と恒常的ホモロジーを統合することで、多スケール恒常的関数、特に多スケール恒常的エントロピーを導入し、バイオ分子構造を特徴付ける。この手法により、タンパク質コンformationの自然なクラスタリングが可能となり、すべてのアルファ、すべてのベータ、および混合タンパク質の正確な分類が達成される。また、トポロジカルエントロピーに基づいて構造の規則性を定量化する新しいタンパク質構造インデックス(PSI)が提案されている。

ABSTRACT

In this paper, we introduce multiscale persistent functions for biomolecular structure characterization. The essential idea is to combine our multiscale rigidity functions with persistent homology analysis, so as to construct a series of multiscale persistent functions, particularly multiscale persistent entropies, for structure characterization. To clarify the fundamental idea of our method, the multiscale persistent entropy model is discussed in great detail. Mathematically, unlike the previous persistent entropy or topological entropy, a special resolution parameter is incorporated into our model. Various scales can be achieved by tuning its value. Physically, our multiscale persistent entropy can be used in conformation entropy evaluation. More specifically, it is found that our method incorporates in it a natural classification scheme. This is achieved through a density filtration of a multiscale rigidity function built from bond and/or dihedral angle distributions. To further validate our model, a systematical comparison with the traditional entropy evaluation model is done. It is found that our model is able to preserve the intrinsic topological features of biomolecular data much better than traditional approaches, particularly for resolutions in the mediate range. Moreover, our method can be successfully used in protein classification. For a test database with around nine hundred proteins, a clear separation between all-alpha and all-beta proteins can be achieved, using only the dihedral and pseudo-bond angle information. Finally, a special protein structure index (PSI) is proposed, for the first time, to describe the "regularity" of protein structures. Essentially, PSI can be used to describe the "regularity" information in any systems.

研究の動機と目的

  • 従来の手法を超えて、複雑なバイオ分子構造を特徴付ける多スケールトポロジカルフレームワークの構築を目的とする。
  • 内在的なトポロジカル特徴を捉えることが難しい準調和的およびカルテシアン座標に基づくエントロピー推定の限界を克服することを目的とする。
  • 恒常的ホモロジーを用いて、自然なクラスタリングおよび構造的規則性の情報をトポロジカル不変量に埋め込むこと。
  • $β_0$ 恒常的エントロピーに基づいて構造的規則性を定量化するタンパク質構造インデックス(PSI)を提案すること。
  • 900タンパク質のデータベースを用いた手法の妥当性評価を行い、中間解像度範囲においてトポロジカル特徴の優れた保存が示されることを目的とする。

提案手法

  • 結合および二面角分布から導出される多スケール剛性関数と密度フィルタリングを組み合わせ、多スケールバーコード表現を生成する。
  • カーネル関数に解像度パラメータを導入することで、複数スケールでの調整が可能となり、多スケール解析が可能となる。
  • 得られたバーコード空間上で、多スケール恒常的関数、特に多スケール恒常的エントロピーを定義する。
  • PSI評価のため、スケールパラメータ $\eta = 5^\circ$ のガウスカーネルを用いてトポロジカルエントロピーを計算する。
  • $\beta_0$、$\beta_1$、$\beta_2$ ベッチ数に対して恒常的ホモロジーを適用し、PSI構築には $\beta_0$ に焦点を当てる。
  • 多スケール剛性関数を密度フィルタリングにより処理し、階層的なトポロジカル特徴を生成する。

実験結果

リサーチクエスチョン

  • RQ1多スケール恒常的関数は、従来のエントロピーモデルに比べ、バイオ分子データの内在的トポロジカル特徴をより良く保持できるか?
  • RQ2提案手法は、トポロジカル構造に基づいてタンパク質コンformationを自然に分類できるか?
  • RQ3$\beta_0$ バーコードから導出される恒常的エントロピーは、タンパク質における構造的規則性の信頼できる指標として機能するか?
  • RQ4解像度パラメータは、中間スケールのバイオ分子特徴におけるトポロジカル特徴の性能にどのように影響するか?
  • RQ5タンパク質構造インデックス(PSI)は、ループや不規則領域を多く含むタンパク質と、安定した二次構造を有するタンパく質を効果的に区別できるか?

主な発見

  • 提案された多スケール恒常的エントロピーモデルは、特に中間解像度範囲において、従来手法よりもバイオ分子データの内在的トポロジカル特徴を顕著に良く保持している。
  • 約900タンパク質のテストデータベースにおいて、二面角および擬似結合角情報のみを用いても、すべてのアルファタンパク質とすべてのベータタンパク質が明確に分離された。
  • アルファおよびベータが混合したタンパク質は、中間範囲の恒常的エントロピー値を示し、重複ケースは僅かにとどまっているため、分類の堅牢性が裏付けられた。
  • タンパク質構造インデックス(PSI)は構造的規則性を効果的に定量化している:ループや不規則領域が多いタンパク質は高いPSI値を示し、アルファヘリックスやベータシートが豊富なタンパク質は低いPSI値を示す。
  • 110タンパク質のテストセットにおいてPSI値は0.000(最も規則的)から2.504(最も不規則)の範囲を示し、構造的組織の感度が確認された。
  • $\beta_1$ および $\beta_2$ バーコードからの高次元トポロジカルエントロピーは、より豊富な情報を提供すると予想されるが、本研究では検討されていない。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。