Skip to main content
QUICK REVIEW

[論文レビュー] A Point Cloud-Based Deep Learning Strategy for Protein-Ligand Binding Affinity Prediction

Yeji Wang, Shuo Wu|arXiv (Cornell University)|Jul 9, 2021
Computational Drug Discovery Methods参考文献 14被引用数 6
ひとこと要約

本研究では、PDBbind-2016データから抽出された3次元点群を用いて、タンパク質-リガンド結合親和性を予測するための新しい深層学習手法を提案する。タンパク質-リガンド複合体を点群として表現するため、PointNetおよびPointTransformerアーキテクチャを適用した結果、それぞれR = 0.831およびR = 0.859のピアソン相関係数を達成し、最先端のモデルと同等の優れた性能を示した。

ABSTRACT

There is great interest to develop artificial intelligence-based protein-ligand affinity models due to their immense applications in drug discovery. In this paper, PointNet and PointTransformer, two pointwise multi-layer perceptrons have been applied for protein-ligand affinity prediction for the first time. Three-dimensional point clouds could be rapidly generated from the data sets in PDBbind-2016, which contain 3 772 and 11 327 individual point clouds derived from the refined or/and general sets, respectively. These point clouds were used to train PointNet or PointTransformer, resulting in protein-ligand affinity prediction models with Pearson correlation coefficients R = 0.831 or 0.859 from the larger point clouds respectively, based on the CASF-2016 benchmark test. The analysis of the parameters suggests that the two deep learning models were capable to learn many interactions between proteins and their ligands, and these key atoms for the interaction could be visualized in point clouds. The protein-ligand interaction features learned by PointTransformer could be further adapted for the XGBoost-based machine learning algorithm, resulting in prediction models with an average Rp of 0.831, which is on par with the state-of-the-art machine learning models based on PDBbind database. These results suggest that point clouds derived from the PDBbind datasets are useful to evaluate the performance of 3D point clouds-centered deep learning algorithms, which could learn critical protein-ligand interactions from natural evolution or medicinal chemistry and have wide applications in studying protein-ligand interactions.

研究の動機と目的

  • タンパク質-リガンド複合体の3次元点群表現を用いた、タンパク質-リガンド結合親和性予測のための深層学習フレームワークの開発。
  • 点群ベースのモデルが、タンパク質-リガンド相互作用の重要な特徴を効果的に捉えられるかの検証。
  • 大規模なPDBbind-2016データセット上で、PointNetとPointTransformerの性能を比較。
  • XGBoostなどの従来の機械学習モデルと、学習された特徴を統合することで予測性能を向上させる手法の探求。

提案手法

  • PDBbind-2016データベースに含まれるタンパク質-リガンド複合体の原子座標を用いて、3次元点群に変換した。
  • PointNetおよびPointTransformerの両方とも、点毎の多層パーセプトロンとしての構造を有し、3次元点群データを処理してエンドツーエンドの特徴抽出を実行した。
  • モデルは、点群に符号化された幾何学的および化学的特徴から、直接的に結合親和性(pKd/pKi)を予測するように学習した。
  • PointTransformerに組み込まれたアテンション機構により、重要な相互作用に寄与する原子に注目できるようになり、特徴表現が向上した。
  • PointTransformerで学習された特徴は、さらにXGBoostに基づく機械学習モデルの入力として用いられ、予測性能の向上が図られた。
  • フレームワークは、最先端の手法と一貫性を持つよう、CASF-2016ベンチマークを用いて評価された。

実験結果

リサーチクエスチョン

  • RQ13次元点群表現は、タンパク質-リガンド相互作用の特徴を結合親和性予測に有効に符号化できるか?
  • RQ2PointNetとPointTransformerは、点群データから意味のある空間的および化学的パターンをどの程度効果的に学習できるか?
  • RQ3PointTransformerに組み込まれたアテンション機構は、タンパク質-リガンド複合体において生物学的に関連する相互作用原子をどの程度明確に特定できるか?
  • RQ4深層学習モデルで学習された特徴は、XGBoostのような従来の勾配ブースティングモデルに効果的に統合可能か?
  • RQ5CASF-2016ベンチマークにおいて、本手法は既存の最先端モデルと比較して、どの程度の予測精度を達成するか?

主な発見

  • PointNetベースのモデルは、CASF-2016ベンチマークでピアソン相関係数R = 0.831を達成し、優れた予測能力を示した。
  • PointTransformerベースのモデルはPointNetを上回り、R = 0.859を達成した。これは、3次元点群からの特徴抽出が優れていることを示している。
  • PointTransformerのアテンション機構により、タンパク質-リガンド相互作用に関与する重要な原子が明確に特定され、解釈可能な特徴可視化が可能になった。
  • XGBoostと組み合わせた場合、PointTransformerの特徴は平均してRp = 0.831の予測性能を示し、最先端のモデルと同等の水準に達した。
  • 本研究では、PDBbindデータから導出された点群が、進化的および医薬的に関連する相互作用を効果的に捉える3次元深層学習モデルの学習に有効であることが確認された。
  • 結果から、点群ベースの深層学習が、創薬におけるタンパク質-リガンド相互作用の研究に有効であることが裏付けられた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。