Skip to main content
QUICK REVIEW

[논문 리뷰] Protein-Nucleic Acid Complex Modeling with Frame Averaging Transformer

Tinglin Huang, Zhenqiao Song|arXiv (Cornell University)|2024. 06. 13.
Bacteriophages and microbial interactionsEnvironmental Science인용 수 3
한 줄 요약

이 논문은 단백질-핵산 복합체를 모델링하기 위해 각 트랜스포머 블록 내부에 프레임 평균화를 통합한 등변성 트랜스포머 아키텍처인 FAFormer을 소개한다. 기하학적으로 인지하는 방식으로 잔기-염기서열 접촉 맵을 예측함으로써, FAFormer은 접촉 맵 예측에서 10퍼센트 이상의 상대적 성능 향상을 달성하고, RoseTTAFoldNA 대비 20~30배 빠른 추론 속도로 비지도 aptamer 필터링을 가능하게 하여 실제 aptamer 데이터셋에서 기존 모델을 능가한다.

ABSTRACT

Nucleic acid-based drugs like aptamers have recently demonstrated great therapeutic potential. However, experimental platforms for aptamer screening are costly, and the scarcity of labeled data presents a challenge for supervised methods to learn protein-aptamer binding. To this end, we develop an unsupervised learning approach based on the predicted pairwise contact map between a protein and a nucleic acid and demonstrate its effectiveness in protein-aptamer binding prediction. Our model is based on FAFormer, a novel equivariant transformer architecture that seamlessly integrates frame averaging (FA) within each transformer block. This integration allows our model to infuse geometric information into node features while preserving the spatial semantics of coordinates, leading to greater expressive power than standard FA models. Our results show that FAFormer outperforms existing equivariant models in contact map prediction across three protein complex datasets, with over 10% relative improvement. Moreover, we curate five real-world protein-aptamer interaction datasets and show that the contact map predicted by FAFormer serves as a strong binding indicator for aptamer screening.

연구 동기 및 목표

  • 단백질-aptamer 결합 예측에서 레이블이 부족한 문제를 해결하기 위해 비지도 학습 접근법을 개발하는 것.
  • 기하학적 딥러닝을 활용해 단백질-핵산 복합체의 접촉 맵 예측을 향상시키는 것.
  • 비용이 많이 드는 실험 데이터가 필요 없이 대규모이고 효율적인 aptamer 필터링을 가능하게 하는 것.
  • 공간적 의미를 유지하면서 트랜스포머 아키텍처에 기하학적 불변성을 통합하는 것.

제안 방법

  • FAFormer는 각 트랜스포머 블록 내부에 포함된 프레임 평균화(FA)를 통해 공간 기하학을 유지하는 새로운 등변성 트랜스포머 아키텍처를 사용한다.
  • 지역 프레임 에지 모듈을 통해 노드와 이웃 간의 국소 쌍방향 상호작용을 기하학적 프레임을 사용해 인코딩한다.
  • 편향된 MLP 어텐션 모듈은 관계 기반 엣지 특징을 어텐션 메커니즘에 통합하여 등변성 좌표 갱신을 가능하게 한다.
  • 글로벌 프레임 FFN 레이어는 글로벌 맥락을 통해 기하 정보를 노드 표현에 융합한다.
  • 모델는 3D 구조에서 잔기-염기서열 접촉 맵을 예측하도록 엔드 투 엔드로 훈련되며, 결합 친화도는 최대 접촉 확률로 추정된다.
  • 추론 속도가 향상되기 위해 MSA에 의존하는 모델 대신 ESMFold가 예측한 비결합 상태의 구조를 사용한다.
Figure 1: (a) The pipeline of contact map prediction between protein and nucleic acid, and applying the predicted results for screening in an unsupervised manner. The affinity score is quantified as the maximum contact probability over all pairs. (b) Comparison between Transformer with vanilla frame
Figure 1: (a) The pipeline of contact map prediction between protein and nucleic acid, and applying the predicted results for screening in an unsupervised manner. The affinity score is quantified as the maximum contact probability over all pairs. (b) Comparison between Transformer with vanilla frame

실험 결과

연구 질문

  • RQ1각 트랜스포머 블록 내부에 통합된 프레임 평균화가 단백질-핵산 복합체의 기하 모델링을 향상시키는가?
  • RQ2제안된 FAFormer 아키텍처가 다양한 단백질 복합체 데이터셋에서 기존의 등변성 모델보다 접촉 맵 예측 성능에서 뛰어나게 되는가?
  • RQ3FAFormer이 예측한 접촉 맵이 비지도 aptamer 필터링에 효과적인 지표가 될 수 있는가?
  • RQ4aptamer 필터링에서 FAFormer은 RoseTTAFoldNA와 같은 대규모 사전학습 모델에 비해 속도와 정확도에서 뛰어나게 되는가?

주요 결과

  • FAFormer은 최신 등변성 모델 대비 세 가지 단백질 복합체 데이터셋에서 접촉 맵 예측 성능이 10퍼센트 이상 향상되었다.
  • 다섯 개인 실제 단백질-aptamer 상호작용 데이터셋에서 FAFormer은 RoseTTAFoldNA를 능가했으며, Top10 및 Top50 정밀도와 PRAUC 점수 모두 높았다.
  • 동일한 필터링 작업에서 FAFormer은 RoseTTAFoldNA 대비 20~30배 빠른 추론 속도를 기록했으며, 평균 추론 시간은 단백질-DNA 기준 32.65초, 단백질-RNA 기준 51.75초였다.
  • 접촉 맵 예측에서 FAFormer은 테스트 세트에서는 RoseTTAFoldNA와 동등한 성능를 보였고, 새로운 타겟에 대한 일반화 능력은 뛰어났다.
  • PDB ID 7DVV 및 7KX9에 대한 케이스 스터디에서, FAFormer이 예측한 접촉 맵은 조밀한 접촉 패턴이 없는 경우에도 실제 값과 밀도 높은 일치를 보였다.
  • 각 트랜스포머 블록 내부에 프레임 평균화를 통합함으로써, 표준 FA 또는 구면 조화 함수 기반 방법보다 더 뛰어난 기하 표현력을 확보했고, 높은 계산 부담 없이도 가능했다.
Figure 2: Overview of FAFormer architecture. The input consists of the node features, coordinates, and edge representations, which are processed by a stack of (b) Biased MLP Attention Module, (c) Local Frame Edge Module, (d) Global Frame FFN, and (e) Gate Function. $\sum$ deontes aggregation, $\cdot
Figure 2: Overview of FAFormer architecture. The input consists of the node features, coordinates, and edge representations, which are processed by a stack of (b) Biased MLP Attention Module, (c) Local Frame Edge Module, (d) Global Frame FFN, and (e) Gate Function. $\sum$ deontes aggregation, $\cdot

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.