Skip to main content
QUICK REVIEW

[論文レビュー] Learning to Predict Mutation Effects of Protein-Protein Interactions by Microenvironment-aware Hierarchical Prompt Learning

Lirong Wu, Yijun Tian|arXiv (Cornell University)|May 16, 2024
Cell Image Analysis Techniques被引用数 4
ひとこと要約

本稿では、事前学習されたプロンプトコードブックを通じてマルチスケール構造的依存関係をモデル化することにより、変異が引き起こすタンパク質-タンパク質相互作用への影響を効率的に予測する、マイクロ環境に配慮した階層的プロンプト学習フレームワーク「Prompt-DDG」を提案する。本手法は、予測性能が最先端水準に達し、事前学習コストを顕著に削減した。

ABSTRACT

Protein-protein bindings play a key role in a variety of fundamental biological processes, and thus predicting the effects of amino acid mutations on protein-protein binding is crucial. To tackle the scarcity of annotated mutation data, pre-training with massive unlabeled data has emerged as a promising solution. However, this process faces a series of challenges: (1) complex higher-order dependencies among multiple (more than paired) structural scales have not yet been fully captured; (2) it is rarely explored how mutations alter the local conformation of the surrounding microenvironment; (3) pre-training is costly, both in data size and computational burden. In this paper, we first construct a hierarchical prompt codebook to record common microenvironmental patterns at different structural scales independently. Then, we develop a novel codebook pre-training task, namely masked microenvironment modeling, to model the joint distribution of each mutation with their residue types, angular statistics, and local conformational changes in the microenvironment. With the constructed prompt codebook, we encode the microenvironment around each mutation into multiple hierarchical prompts and combine them to flexibly provide information to wild-type and mutated protein complexes about their microenvironmental differences. Such a hierarchical prompt learning framework has demonstrated superior performance and training efficiency over state-of-the-art pre-training-based methods in mutation effect prediction and a case study of optimizing human antibodies against SARS-CoV-2.

研究の動機と目的

  • タンパク質-タンパク質相互作用研究におけるアノテート済みの実験的変異データの不足に取り組む。
  • 複数の構造的スケールにわたる複雑な高次依存関係をモデル化できない既存の事前学習手法の限界を克服する。
  • 変異が引き起こす局所的マイクロ環境の変化を、変異後の複合体構造を明示的に予測することなく捉える。
  • ΔΔG予測の高い予測性能を維持しつつ、事前学習の計算コストを低減する。
  • 特にSARS-CoV-2中和抗体に対して、効果的なバーチャルスクリーニングと機能的部位の同定を可能にする。

提案手法

  • 異なる構造的スケール(例:アミノ酸、二面角、局所的コンformation)における共通するマイクロ環境パターンを独立に符号化する階層的プロンプトコードブックを構築する。
  • 各変異周辺のアミノ酸種別、角度統計、局所的コンformational 変化を統合的にモデル化する、マスクされたマイクロ環境モデリング事前学習タスクを設計する。
  • コードブックからの階層的プロンプトを組み合わせることで、それぞれの変異に対するマイクロ環境に配慮したプロンプトを生成し、相同型と変異型複合体の差異を捉える。
  • 学習済みプロンプトをコンテキスト内表現として用い、下流のΔΔG予測をガイドする。これにより、高コストなエンドツーエンドのファインチューニングを回避する。
  • プロンプトコードブックを活用することで、各新しい変異に対してモデルを再訓練することなく、柔軟かつマルチスケールの構造的文脈を符号化可能にする。
  • 相同型と変異型マイクロ環境表現を一致させるコントラスト学習目的関数を用いて、フレームワークをエンドツーエンドで学習する。
Figure 1: Comparison of our Prompt-DDG with three state-of-the-art methods in effectiveness (per-structure Pearson and Spearman) and training efficiency (for pre-training and $\Delta\Delta G$ prediction). where Prompt-DDG outperforms the other methods a lot in both effectiveness and efficiency, espe
Figure 1: Comparison of our Prompt-DDG with three state-of-the-art methods in effectiveness (per-structure Pearson and Spearman) and training efficiency (for pre-training and $\Delta\Delta G$ prediction). where Prompt-DDG outperforms the other methods a lot in both effectiveness and efficiency, espe

実験結果

リサーチクエスチョン

  • RQ1階層的プロンプトコードブックは、変異周辺のタンパク質マイクロ環境におけるマルチスケール構造的依存関係を効果的にモデル化できるか?
  • RQ2マスクされたマイクロ環境モデリング事前学習タスクは、標準的な事前学習タスクと比較して、下流のΔΔG予測性能を向上させるか?
  • RQ3プロンプトベースの表現学習は、エンドツーエンドの事前学習と比較して、ΔΔG予測において優れた精度と効率性を達成できるか?
  • RQ4本モデルは、SARS-CoV-2の抗体最適化という実世界の応用にどの程度一般化可能か?
  • RQ5予測精度と学習効率の両面で、最先端手法と比較して本モデルの性能はどの程度か?

主な発見

  • Prompt-DDGは、SKEMPI v2.0ベンチマークにおいて、1構造あたりのピアソン相関係数(0.6557)とスピアマン相関係数(0.5691)が最高を記録し、すべてのベースラインを上回った。
  • RDE-Network や DiffAffinity と比較して、事前学習時間を50%以上短縮したが、優れた性能を維持した。
  • SARS-CoV-2抗体最適化の事例研究において、Prompt-DDGは既知の5つの好ましい変異を上位40%内にランク付けし、平均ランク10.69%を達成した。これは、RDE-Network(18.26%)や DiffAffinity(24.49%)を顕著に上回った。
  • 唯一、Prompt-DDGが上位10%内に5つのうち4つの主要な中和変異を正しく同定した。これは、強力な一般化能力と精度を示している。
  • アブレーションスタディの結果、マスクされたマイクロ環境モデリングタスクにおける最適なマスク率は0.10であり、ピアソン相関とスピアマン相関の両方で最高の性能を発揮した。
  • 階層的プロンプトコードブックにより、マルチスケールのマイクロ環境特徴の有効なモデル化が可能となり、変異効果の一般化と解釈可能性が向上した。
Figure 2: Left: A high-level overview of microenvironment-aware hierarchical prompt learning and adaptation framework for efficient $\Delta\Delta G$ prediction (Prompt-DDG). Right: Illustration of a hierarchical pre-training task by Masked Microenvironment Modeling (MMM).
Figure 2: Left: A high-level overview of microenvironment-aware hierarchical prompt learning and adaptation framework for efficient $\Delta\Delta G$ prediction (Prompt-DDG). Right: Illustration of a hierarchical pre-training task by Masked Microenvironment Modeling (MMM).

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。