Skip to main content
QUICK REVIEW

[論文レビュー] Fine-tuning Protein Language Models with Deep Mutational Scanning improves Variant Effect Prediction

Aleix Lafita, Ferran Gonzalez Hernandez|arXiv (Cornell University)|May 10, 2024
Machine Learning in Bioinformatics被引用数 10
ひとこと要約

論文は Normalised Log-odds Ratio (NLR) を導入する。DMSデータで訓練されたProtein Language Models (PLMs) の軽量なファインチューニングヘッドで、ベンチマーク全体に渡るミスセンス変異の影響予測を改善する。

ABSTRACT

Protein Language Models (PLMs) have emerged as performant and scalable tools for predicting the functional impact and clinical significance of protein-coding variants, but they still lag experimental accuracy. Here, we present a novel fine-tuning approach to improve the performance of PLMs with experimental maps of variant effects from Deep Mutational Scanning (DMS) assays using a Normalised Log-odds Ratio (NLR) head. We find consistent improvements in a held-out protein test set, and on independent DMS and clinical variant annotation benchmarks from ProteinGym and ClinVar. These findings demonstrate that DMS is a promising source of sequence diversity and supervised training data for improving the performance of PLMs for variant effect prediction.

研究の動機と目的

  • ミスセンス変異の機能的影響予測をゼロショットPLM性能超えへ向上させる動機付け。
  • 複数アッセイからのDMSスコアを活用する正規化とファインチューニングのパイプラインを提案。
  • 保持分割されたタンパク質と独立したベンチマークでの改善を実証。
  • トレーニングデータが限定的な場合の一般化を評価し、タンパク質間でのモデル性能を分析。
  • より多くのDMSデータとMSAベースのPLMを統合する将来の方向性とスケーラビリティを議論。

提案手法

  • アッセイ間でDMSスコアを正規化し、同義語の平均を0、ナンセンスの平均を-1にするようリスケールし、[-2, 2]へクリップ。
  • Normalised Log-odds Ratio (NLR) ヘッドを導入し、野生型配列ごとに全置換のログオッズ比の行列を計算。
  • NLRヘッドとともにESM-1vエンコーダをファインチューニング(ESM-1b/ESM-2と比較)し、推論時に5つのモデルチェックポイントの予測を平均。
  • 25タンパク質(109,215変異)によるDMSデータを用い、5-foldクロスバリデーションとその後のフルトレーニングで訓練。
  • 保持分割されたMaveDBテストタンパク質、ProteinGym DMSアッセイ、およびClinVarの致病性/正常変異で評価;タンパク質ごとの性能とベースラインのゼロショット結果を分析。
Figure 1: Methods overview. A) Preparation of normalised DMS functional scores from a subset of MaveDB experiments. The mean scores of synonymous and nonsense variants are used to create a common scale across assays and proteins. B) Fine-tuning pipeline for ESM-1v models using the Normalised Log-odd
Figure 1: Methods overview. A) Preparation of normalised DMS functional scores from a subset of MaveDB experiments. The mean scores of synonymous and nonsense variants are used to create a common scale across assays and proteins. B) Fine-tuning pipeline for ESM-1v models using the Normalised Log-odd

実験結果

リサーチクエスチョン

  • RQ1NLRファインチューニングはゼロショット性能を超えたPLMベースの変異影響予測を改善できるか?
  • RQ2共通スケールで多様なDMSデータを統合することは独立したベンチマークでの予測を向上させるか?
  • RQ3NLRは異なるPLMアーキテクチャ(ESM-1v、ESM-1b、ESM-2)でどのように性能へ影響し、トレーニングデータ量は利得にどのように影響するか?
  • RQ4基準となるゼロショット精度が低いタンパク質や前訓練データで表現が乏しいタンパク質で性能が向上するか?

主な発見

  • NLRファインチューニングはMaveDBテストタンパク質のマイクロ平均Spearman相関を0.478から0.503へ改善(+5.2%)。
  • NLRファインチューニングはProteinGymの平均Spearman相関を0.331から0.396へ改善(+19.6%)。
  • NLRファインチューニングはClinVar auROCを0.891から0.902へ改善(+1.23%)。
  • タンパク質ごとのClinVar分析は、特に基礎のauROCが低いタンパク質で一貫した改善を示す。
  • ESM-1bおよびESM-2もNLRファインチューニングの恩恵を受け、いくつかの設定でProteinGymに最大25.6%の相対Spearman利得を達成;単一アーキテクチャを超えた利得。
  • DMSデータが増えるほど改善が拡大し、前訓練データに表現されていないタンパク質(例: ウイルスタンパク質)での利得がわずかに大きい。
Figure 2: Results after NLR fine-tuning of ESM-1v models across benchmarks. A) Performance in the five MaveDB test proteins. ProteinGym DMS assays and ClinVar pathogenic variants. B) Spearman correlation in MaveDB test proteins. Mean $\pm$ standard deviation (std) of 50 bootstrapped samples. C) Spea
Figure 2: Results after NLR fine-tuning of ESM-1v models across benchmarks. A) Performance in the five MaveDB test proteins. ProteinGym DMS assays and ClinVar pathogenic variants. B) Spearman correlation in MaveDB test proteins. Mean $\pm$ standard deviation (std) of 50 bootstrapped samples. C) Spea

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。