Skip to main content
QUICK REVIEW

[論文レビュー] Persistent cohomology for data with multicomponent heterogeneous information

Zixuan Cang, Guo‐Wei Wei|arXiv (Cornell University)|Jul 29, 2018
Topological and Geometric Data Analysis参考文献 41被引用数 7
ひとこと要約

本稿では、原子電荷や静電ポテンシャルなどのマルチコンponentで異種のデータを幾何学的構造とともにより一貫したトポロジカル不変量に統合する、持続的コhomologyフレームワークを提案する。単体複体上で滑らかなコサイクルを計算することにより、標準的なパーシステントバークォートに物理的情報を追加し、特に静電気が組み込まれた場合に、タンパク質-リガンド結合親和定数の予測精度を顕著に向上させる。

ABSTRACT

Persistent homology is a powerful tool for characterizing the topology of a data set at various geometric scales. When applied to the description of molecular structures, persistent homology can capture the multiscale geometric features and reveal certain interaction patterns in terms of topological invariants. However, in addition to the geometric information, there is a wide variety of non-geometric information of molecular structures, such as element types, atomic partial charges, atomic pairwise interactions, and electrostatic potential function, that is not described by persistent homology. Although element specific homology and electrostatic persistent homology can encode some non-geometric information into geometry based topological invariants, it is desirable to have a mathematical framework to systematically embed both geometric and non-geometric information, i.e., multicomponent heterogeneous information, into unified topological descriptions. To this end, we propose a mathematical framework based on persistent cohomology. In our framework, non-geometric information can be either distributed globally or resided locally on the datasets in the geometric sense and can be properly defined on topological spaces, i.e., simplicial complexes. Using the proposed persistent cohomology based framework, enriched barcodes are extracted from datasets to represent heterogeneous information. We consider a variety of datasets to validate the present formulation and illustrate the usefulness of the proposed persistent cohomology. It is found that the proposed framework using cohomology boosts the performance of persistent homology based methods in the protein-ligand binding affinity prediction on massive biomolecular datasets.

研究の動機と目的

  • パーサイステントホモロジーが原子電荷や静電ポテンシャルなどの非幾何的分子性質を捉えることの限界を是正すること。
  • 幾何的および非幾何的情報をトポロジカルデータ解析において統合する数学的フレームワークを構築すること。
  • 元素種別、部分電荷、相互作用エネルギーなどの異種データを、トポロジカル記述子に体系的に埋め込むこと。
  • 特にタンパク質-リガンド結合親和定数の予測性能を向上させることを目的とした、バイオ分子モデリングにおけるトポロジカル手法の予測力強化。

提案手法

  • 距離またはスケールパラメータに基づくフィルトレーションを用いて、点群データから単体複体を構築する。
  • 単体複体上に重み付きグラフラプラシアンを定義し、非幾何的情報を表す滑らかなコサイクルを計算する。
  • 滑らかなコサイクルを単体上の関数として用い、標準的なパーシステントバークォートに物理的情報を追加する。
  • 幾何的および非幾何的情報を両方含む enriched バークォートを比較するため、修正されたワサーティン距離を導入する。
  • 頂点に静電ポテンシャル値を割り当て、コホモロジーを介してそれを伝搬させることで、バイオ分子データセットにフレームワークを適用する。
  • グリッドサーチによるハイパーパramータチューニングを経た勾配ブースティングを用い、enriched バークォートから結合親和定数を予測する。

実験結果

リサーチクエスチョン

  • RQ1パーサイステントコホモロジーを用いて、原子部分電荷などの非幾何的分子性質をトポロジカル不変量に埋め込むことができるか?
  • RQ2物理的情報をパーシステントバークォートに追加することで、タンパク質-リガンド結合親和定数予測における機械学習モデルの性能にどのような影響を与えるか?
  • RQ3ホモロジーと比較して、コホモロジーに基づく記述子は、ループや空洞といったトポロジカル特徴と物理的性質をより良く局在化・関連付けることができるか?
  • RQ4大規模なバイオ分子データセットにおいて、静電情報を組み込むことで、トポロジカルモデルの予測精度がどの程度向上するか?
  • RQ5提案されたフレームワークは、ファンデルワールス相互作用や原子相互作用エネルギーなどの他の物理的性質に対しても一般化可能か?

主な発見

  • 持続的コホモロジーフレームワークは、単体複体上の滑らかなコサイクルを介して、静電ポテンシャルなどの非幾何的情報をトポロジカル不変量に成功して埋め込んでいる。
  • バークォートに静電情報を組み込むことで、テストしたPDBbind全バージョンでタンパク質-リガンド結合親和定数の予測性能が向上し、特にv2016で最大の向上が観察された(ピアソン相関係数:0.778 対 0.767、標準的パーサイステントホモロジー)。
  • 静電情報を組み込んだ持続的コホモロジーを用いることで、PDBbind v2016コアセットにおいて、pKdの中央値ピアソン相関係数が0.833に達し、近似的最適な性能を達成した。
  • 修正されたワサーティン距離は、異種データを含むenriched バークォート間の類似度を効果的に定量化し、トポロジカル記述子の堅牢な比較を可能にした。
  • 本手法は、0次およびそれ以上の次元の持続的コホモロジーにおいて、標準的パーサイステントホモロジーを常に上回る性能を示し、特に物理的相互作用に関連する生物学的に意味のある特徴を捉える点で優れた効果を示した。
  • 本手法は静電気を超えて一般化可能であり、ファンデルワールス相互作用や原子相互作用エネルギーなどの他の物理的性質へも拡張可能であり、複雑なデータセットのより豊かなトポロジカルモデリングを可能にする。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。