Skip to main content
QUICK REVIEW

[論文レビュー] Embracing assay heterogeneity with neural processes for markedly improved bioactivity predictions

Lucian Chan, Marcel L. Verdonk|arXiv (Cornell University)|Aug 17, 2023
Computational Drug Discovery MethodsComputer Science被引用数 3
ひとこと要約

この論文では、メタラーニングフレームワークであるMetaBindを提案する。このフレームワークは、神経過程を用いてバイオアクティビティデータ内のアッセイの異質性をモデル化し、多様なタンパク質標的およびアッセイタイプにわたって、非常に高い精度と適応性を持つアフィニティ予測を可能にする。スパースで異質なデータから学習し、タスク固有のサポートセットを活用することで、従来のモデルと比較して著しく低い誤差率を達成し、最先端の性能を発揮する。

ABSTRACT

Predicting the bioactivity of a ligand is one of the hardest and most important challenges in computer-aided drug discovery. Despite years of data collection and curation efforts by research organizations worldwide, bioactivity data remains sparse and heterogeneous, thus hampering efforts to build predictive models that are accurate, transferable and robust. The intrinsic variability of the experimental data is further compounded by data aggregation practices that neglect heterogeneity to overcome sparsity. Here we discuss the limitations of these practices and present a hierarchical meta-learning framework that exploits the information synergy across disparate assays by successfully accounting for assay heterogeneity. We show that the model achieves a drastic improvement in affinity prediction across diverse protein targets and assay types compared to conventional baselines. It can quickly adapt to new target contexts using very few observations, thus enabling large-scale virtual screening in early-phase drug discovery.

研究の動機と目的

  • ドラッグディスカバリにおけるスパースで異質なバイオアクティビティデータの課題に対処し、既存の予測モデルの性能と移行性を制限する要因を解消すること。
  • 実験的条件やアッセイタイプの違いがある中でも、異なるアッセイ間の情報の相乗効果を効果的に活用できるモデルを開発すること。
  • 新しい標的やアッセイに、わずかな測定済みの活性データのみで迅速に適応できることを可能にし、大規模なバーチャルスクリーニングを支援すること。
  • アッセイレベルのばらつきを明示的にモデル化することで、ノイズとして扱うのではなく、リガンドアフィニティ予測の精度と頑健性を向上させること。

提案手法

  • 各アッセイを別個の学習タスクとして扱い、K個の共分散クラスタにグループ化することで、共通の構造的パターンをモデル化する階層的メタラーニングフレームワークを採用する。
  • 神経過程アーキテクチャを用いて、観測済みのリガンド-タンパク質ペアとその測定バイオアクティビティの小さなサポートセットに条件付けられた関数の分布を学習する。
  • リガンド表現にはグラフ畳み込みネットワークを、タンパク質埋め込みには畳み込みネットワークを用い、タスク間で情報を統合するためにクロス共分散演算子を導入する。
  • トレーニング中にKullback-Leiblerダイバージェンス損失を適用し、各アッセイの完全な共分散構造を、その観測値のランダムサブセットから再構成できるように促進する。
  • フレームワークは、ChEMBL 30データ上でエンドツーエンドにトレーニングされ、キニ、Kd、IC50、EC50のエンドポイントを持つ高信頼性のヒト由来阻害アッセイに絞り込まれている。
  • モデルの評価にはタスクレベルの指標(T-RMSE、T-MAE)とペアドアッセイスプリットを用い、同じ標的を標的にするが異なる構造-活性関係を示すアッセイ間の差を識別できる能力を評価する。
Figure 1: Illustration of the MetaBind approach. The model constructs an assay-specific SAR using a local support set of per-assay observations to predict affinities for unobserved protein-ligand pairs. In meta-learning terms, each assay is thus treated as a task, and assays are clustered into a sma
Figure 1: Illustration of the MetaBind approach. The model constructs an assay-specific SAR using a local support set of per-assay observations to predict affinities for unobserved protein-ligand pairs. In meta-learning terms, each assay is thus treated as a task, and assays are clustered into a sma

実験結果

リサーチクエスチョン

  • RQ1メタラーニングモデルは、多様なタンパク質標的において、アッセイの異質性を効果的に捉え、それを活用してバイオアクティビティ予測を改善できるか?
  • RQ2少量のラベル付きデータで、新しい未観測の標的やアッセイにどの程度一般化できるか?
  • RQ3生物学的に類似しているが、実験的に異質なアッセイをどの程度正確に区別できるか?
  • RQ4アッセイレベルのばらつきを考慮することで、従来のモデルと比較して予測精度が顕著に向上するか?

主な発見

  • MetaBindはテストセットでタスクレベルの平均二乗誤差(T-RMSE)を0.48 pK単位に達成し、従来のベースラインと比べて著しく改善された。
  • モデルは強力なゼロショット一般化性能を示し、178の未観測の標的タンパク質においてT-MAEが0.36 pK単位に達した。トレーニングとテストセット間のリガンド類似度中央値は0.32であった。
  • ペアドアッセイスプリットにおいて、同じタンパク質標的を標的にするが、異なる構造-活性関係を示すアッセイ間の差を効果的に特定・モデル化できており、異質性への頑健な対処能力を示している。
  • 生化学的および細胞ベースのアッセイを含む多様なアッセイタイプにおいて、標準的なディープラーニングおよびガウス過程ベースラインを上回る性能を発揮した。
  • Kullback-Leibler損失項の使用により、部分観測セットからの共分散構造の効果的再構成が可能となり、一般化性能と不確実性のキャリブレーションが向上した。
Figure 2: Observed and predicted heterogeneities in structure-activity relationships. (a-c) Correlation of the bioactivity read-out for congeneric series of three different assay pairs. The pairs (A,B) and (C,D) share the same protein variants, but differ in assay type and conditions. The assays of
Figure 2: Observed and predicted heterogeneities in structure-activity relationships. (a-c) Correlation of the bioactivity read-out for congeneric series of three different assay pairs. The pairs (A,B) and (C,D) share the same protein variants, but differ in assay type and conditions. The assays of

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。