Skip to main content
QUICK REVIEW

[논문 리뷰] Improved motif-scaffolding with SE(3) flow matching

Jason Yim, Andrew M. Campbell|PubMed|2024. 01. 08.
Protein Structure and Dynamics참고 문헌 15인용 수 8
한 줄 요약

본 논문은 FrameFlow에 모티프 조건화된 암모타이제이션과 모티프 가이드를 도입하여 모티프-스캐폴딩을 수행하고, RFdiffusion에 비해 설계가능성은 유지하면서 골격 다양성을 더 높게 달성한다.

ABSTRACT

Protein design often begins with the knowledge of a desired function from a motif which motif-scaffolding aims to construct a functional protein around. Recently, generative models have achieved breakthrough success in designing scaffolds for a range of motifs. However, generated scaffolds tend to lack structural diversity, which can hinder success in wet-lab validation. In this work, we extend FrameFlow, an SE(3) flow matching model for protein backbone generation, to perform motif-scaffolding with two complementary approaches. The first is motif amortization, in which FrameFlow is trained with the motif as input using a data augmentation strategy. The second is motif guidance, which performs scaffolding using an estimate of the conditional score from FrameFlow without additional training. On a benchmark of 24 biologically meaningful motifs, we show our method achieves 2.5 times more designable and unique motif-scaffolds compared to state-of-the-art. Code: https://github.com/microsoft/protein-frame-flow.

연구 동기 및 목표

  • 모티프-스캐폴딩을 확장된 SE(3) 흐름 매칭으로 두 가지 전략(모티프 암모타이제이션 및 모티프 가이드)로 개선한다.
  • PDB 기반 모티프-스캐폴딩 벤치마크에서 조건부(암모타이즈된)와 무조건부(가이드) 접근법을 비교한다.
  • FrameFlow 변형들이 RFdiffusion보다 동등하거나 더 나은 설계가능성과 더 큰 골격 다양성을 달성함을 보인다.
  • 가벼운 FrameFlow 모델이 현재의 최첨단 방법들보다 더 적은 파라미터와 학습 자원을 필요로 한다는 것을 Demonstrate 한다.

제안 방법

  • FrameFlow에 모티프 조건화를 추가하여 주어진 모티프 주위에 골격을 생성한다(모티프 암모타이제이션).
  • 추가 학습 없이 무조건부 FrameFlow 모델을 사용하여 모티프로 샘플링 궤적을 조건화해 모티프 가이드를 개발한다.
  • 변환 벡터 필드를 분해한 벡터 필드를 이용한 Riemannian 흐름 매칭으로 SE(3) 백본 표현을 모델링한다(SO(3) 회전 및 평행이동에 대한).
  • 비라벨링 PDB에서 모티프 분포를 시뮬레이션하기 위해 모티프 데이터 증강으로 FrameFlow를 모티프-암모타이즈드 방식으로 학습한다.
  • 생성에 500 타임스텝의 Euler-Maruyama 샘플링을 사용하고 설계가능성 및 다양성 지표로 평가한다.
  • 모티프-스캐폴딩 벤치마크에서 RFdiffusion과 TDS와 비교한다.
Figure 1: We present two strategies for motif-scaffolding. Top : motif amortization trains a flow model to condition on the motif (blue) and generate the scaffold (red). During training, only the scaffold is corrupted with noise. Bottom : motif guidance re-purposes a flow model that is trained to ge
Figure 1: We present two strategies for motif-scaffolding. Top : motif amortization trains a flow model to condition on the motif (blue) and generate the scaffold (red). During training, only the scaffold is corrupted with noise. Bottom : motif guidance re-purposes a flow model that is trained to ge

실험 결과

연구 질문

  • RQ1모티프 암모타이제이션 또는 모티프 가이드가 SE(3) 흐름 매칭에 대해 이전의 최첨단(RFdiffusion)보다 모티프-스캐폴딩 성능을 개선할 수 있는가?
  • RQ2조건부(암모타이즈)와 무조건부(가이드) 접근법은 설계가능성과 골격 다양성에 차이가 있는가?
  • RQ3FrameFlow가 더 가볍고 학습 가능하면서 모티프마다 더 많고 설계가능한 골격을 제공하는가?
  • RQ4무조건부 백본 결과가 FrameFlow-가이드를 통한 모티프-스캐폴딩의 신뢰성을 뒷받침하는가?

주요 결과

  • 모티프 암모타이제이션을 활용한 FrameFlow는 벤치마크에서 21개의 모티프를 해결했고 RFdiffusion은 20개를 달성했다.
  • FrameFlow-가이드는 모티프를 20개 해결하여 RFdiffusion의 성능과 일치한다.
  • 모티프 전체에 걸쳐 FrameFlow 방식은 RFdiffusion보다 고유한 설계가능 골격을 2.5배 더 많이 생성한다.
  • 무조건부 FrameFlow는 RFdiffusion과 비슷한 설계가능성을 제공하면서 더 높은 다양성과 새로움을 달성한다.
  • FrameFlow 조건화 접근은 네트워크를 3배 작게(make 16.8M vs 59.8M 파라미터) 하고 사전 학습이 필요 없다.
  • FrameFlow를 이용한 모티프-스캐폴딩은 RFdiffusion 및 TDS보다 더 큰 골격 다양성을 달성한다.
Figure 2: Motif data augmentation. Each protein in the dataset does not come with pre-defined motif-scaffold annotations. Instead, we construct plausible motifs at random to simulate sampling from the distribution of motifs and scaffolds.
Figure 2: Motif data augmentation. Each protein in the dataset does not come with pre-defined motif-scaffold annotations. Instead, we construct plausible motifs at random to simulate sampling from the distribution of motifs and scaffolds.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.