Skip to main content
QUICK REVIEW

[論文レビュー] Improving Protein Optimization with Smoothed Fitness Landscapes

Andrew Kirjner, Jason Yim|arXiv (Cornell University)|Jul 2, 2023
Genomics and Phylogenetic StudiesBiochemistry, Genetics and Molecular Biology被引用数 3
ひとこと要約

本稿では、グラフ信号処理とTikunov正則化を用いてノイズの多いフィットネスランドスケープを平滑化することで、タンパク質最適化を向上させるGibbsサンプリングとグラフベース平滑化(GGS)を提案する。GGSは、インシリコ評価においてトレーニングセット比で2.5倍のフィットネス向上を達成し、GFPおよびAAVベンチマークにおいて、平滑化されたエネルギーに基づくモデルと勾配誘導型Gibbsサンプリングを組み合わせることで、先行手法を上回る性能を示した。

ABSTRACT

The ability to engineer novel proteins with higher fitness for a desired property would be revolutionary for biotechnology and medicine. Modeling the combinatorially large space of sequences is infeasible; prior methods often constrain optimization to a small mutational radius, but this drastically limits the design space. Instead of heuristics, we propose smoothing the fitness landscape to facilitate protein optimization. First, we formulate protein fitness as a graph signal then use Tikunov regularization to smooth the fitness landscape. We find optimizing in this smoothed landscape leads to improved performance across multiple methods in the GFP and AAV benchmarks. Second, we achieve state-of-the-art results utilizing discrete energy-based models and MCMC in the smoothed landscape. Our method, called Gibbs sampling with Graph-based Smoothing (GGS), demonstrates a unique ability to achieve 2.5 fold fitness improvement (with in-silico evaluation) over its training set. GGS demonstrates potential to optimize proteins in the limited data regime. Code: https://github.com/kirjner/GGS

研究の動機と目的

  • 限られた実験データのもとで、高次元的でノイズが多く、スパースなフィットネスランドスケープにおけるタンパク質最適化の課題に対処すること。
  • ノイズを低減し、局所最適解を回避するために、フィットネスランドスケープの平滑化を通じてタンパク質設計を改善すること。
  • グラフベース正則化を用いて、限られたデータ環境下でも効果的な最適化を可能にする手法を開発すること。
  • GFPおよびAAVタンパク質最適化ベンチマークにおいて、平滑化と勾配ベースのサンプリングを組み合わせた新規な手法を用いて、最先端の性能を示すこと。

提案手法

  • ノードをタンパク質配列、エッジを類似性として定義するグラフとしてタンパク質配列を定式化し、フィットネス値をノード属性として扱う。
  • グラフラプラシアンを用いてTikunov正則化を適用し、フィットネスランドスケープを平滑化することで、ノイズを低減し、信号を強化する。
  • 平滑化されたフィットネスランドスケープ上でニューラルネットワークを訓練し、最適化のための離散的エネルギーに基づくモデルを構築する。
  • 勾配を活用して変異を段階的にサンプリングするGibbs With Gradients(GWG)を用い、フィットネス向上に寄与する変更を優遇する。
  • 勾配誘導型の変異提案を繰り返し適用し、高いフィットネスを持つ配列へ向かう軌道をガイドする。
  • グラフサイズ、平滑化重みγ、サンプリングラウンド数といったハイパーパrameterを最適化し、性能と収束性のバランスを取る。
Figure 1: Overview. (A) Protein optimization is challenging due to a noisy fitness landscape where the starting dataset (unblurred) is a fraction of the landscape with the highest fitness sequences hidden (blurred). (B) We develop Graph-based Smoothing (GS) to estimate a smoothed fitness landscape f
Figure 1: Overview. (A) Protein optimization is challenging due to a noisy fitness landscape where the starting dataset (unblurred) is a fraction of the landscape with the highest fitness sequences hidden (blurred). (B) We develop Graph-based Smoothing (GS) to estimate a smoothed fitness landscape f

実験結果

リサーチクエスチョン

  • RQ1ノイズが多く、データが限られた状況下でも、タンパク質のフィットネスランドスケープに対するグラフベースの平滑化は、最適化性能の向上に寄与するか?
  • RQ2フィットネスランドスケープの平滑化は、タンパク質設計における機械学習モデルの一般化能力と予測精度にどのように影響するか?
  • RQ3平滑化されたエネルギーに基づくモデルと勾配誘導型サンプリング(GWG)を組み合わせることで、ベースライン手法と比較してどの程度フィットネス向上が達成されるか?
  • RQ4グラフサイズや平滑化強度といったハイパーパrameterのうち、最適化性能に最も顕著に影響を与えるものは何か?
  • RQ5実世界のタンパク質ベンチマーク(GFPおよびAAV)において、限られたデータ制約のもとで、提案手法が最先端の結果を達成できるか?

主な発見

  • インシリコ評価において、GGSはトレーニングセット比で2.5倍の予測フィットネス向上を達成し、優れた外挿能力を示した。
  • 平滑化は一般化性能を顕著に向上させた:ハードフィルタリングを施した4%のデータのみで学習したにもかかわらず、スコアレーティング評価者との相関が高く維持された。
  • 最適なグラフサイズは250,000ノードであった。これは近似精度と計算コストのバランスを取る上で最適であった。
  • 過剰な平滑化(γ = 10.0)はAAVにおいて性能を低下させた。これは、γがフィットネスランドスケープごとに慎重にチューニングされる必要があることを示唆している。
  • 滑らかなランドスケープ下では15ラウンド以内に収束が達成され、GFPおよびAAVの両タスクで温度τ = 0.1が最良の性能を示した。
  • 平滑化はGGSにとどまらず、ベースライン手法に対しても性能向上をもたらした。一部の手法では最大3倍の向上が観察され、ランドスケープ平滑化の広範な有用性が裏付けられた。
Figure 2: Steps in graph-based smoothing on proteins illustrated with a fictitious data of length 2 sequences with vocabulary $\{A,B\}$ . Above each node are corresponding fitness values. Solid nodes are those in our training set while dashed nodes are augmented via point mutations to increase the s
Figure 2: Steps in graph-based smoothing on proteins illustrated with a fictitious data of length 2 sequences with vocabulary $\{A,B\}$ . Above each node are corresponding fitness values. Solid nodes are those in our training set while dashed nodes are augmented via point mutations to increase the s

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。