Skip to main content
QUICK REVIEW

[논문 리뷰] ProtST: Multi-Modality Learning of Protein Sequences and Biomedical Texts

Minghao Xu, Xinyu Yuan|arXiv (Cornell University)|2023. 01. 28.
Machine Learning in Bioinformatics인용 수 30
한 줄 요약

ProtST는 ProtDescribe를 도입한다, 페어드 단백질 시퀀스-텍스트 데이터셋과 단백질 시퀀스와 생의학 텍스트를 정렬하는 다중 모달 사전학습 프레임워크를 통해 단백질 표현을 향상시키고 제로샷 예측 및 텍스트-단백질 검색을 가능하게 한다.

ABSTRACT

Current protein language models (PLMs) learn protein representations mainly based on their sequences, thereby well capturing co-evolutionary information, but they are unable to explicitly acquire protein functions, which is the end goal of protein representation learning. Fortunately, for many proteins, their textual property descriptions are available, where their various functions are also described. Motivated by this fact, we first build the ProtDescribe dataset to augment protein sequences with text descriptions of their functions and other important properties. Based on this dataset, we propose the ProtST framework to enhance Protein Sequence pre-training and understanding by biomedical Texts. During pre-training, we design three types of tasks, i.e., unimodal mask prediction, multimodal representation alignment and multimodal mask prediction, to enhance a PLM with protein property information with different granularities and, at the same time, preserve the PLM's original representation power. On downstream tasks, ProtST enables both supervised learning and zero-shot prediction. We verify the superiority of ProtST-induced PLMs over previous ones on diverse representation learning benchmarks. Under the zero-shot setting, we show the effectiveness of ProtST on zero-shot protein classification, and ProtST also enables functional protein retrieval from a large-scale database without any function annotation.

연구 동기 및 목표

  • 생의학 텍스트에 기술된 단백질 특성으로 단백질 시퀀스 표현을 보강한다.
  • ProtDescribe를 만들어 시퀀스를 풍부한 속성 설명과 짝지운다.
  • 시퀀스 모델링 능력을 보존하면서 속성 정보를 주입하는 다중 모달 사전학습 과제를 개발한다.
  • 이후 감독학습 및 제로샷 단백질 분류와 검색을 가능하게 한다.

제안 방법

  • 단일 모달 마스크 예측을 사용하여 단백질 시퀀스의 공동진화 신호를 보존한다.
  • 시퀀스와 텍스트 표현 간의 대조적 InfoNCE 손실로 다중 모달 표현 정렬을 적용한다.
  • 잔류물과 텍스트 토큰 간의 상호의존성을 모델링하는 융합 모듈을 활용한 다중 모달 마스크 예측을 도입한다.
  • 고정된 생의학 언어 모델(PubMedBERT)과 융합 모듈을 활용하여 PLM(ProtBert/ESM 변형)을 사전 학습한다.
  • 레이블 설명 및 제로샷 텍스트-단백질 검색을 사용한 제로샷 분류를 지원하기 위해 표현들을 정렬한다.

실험 결과

연구 질문

  • RQ1생의학 텍스트로 기술된 단백질 특성 설명이 시퀀스 전용 PLM을 넘는 단백질 시퀀스 표현을 향상시킬 수 있는가?
  • RQ2다중 모달 정렬 및 교차 모달 마스킹이 다운스트림의 단백질 위치, 적합성, 기능 예측을 어느 정도 향상시키는가?
  • RQ3ProtST로 유도된 모델이 기능 주석 없이 레이블 설명과 제로샷 검색을 사용한 효과적인 제로샷 단백질 분류를 가능하게 하는가?
  • RQ4Swiss-Prot 기반인 ProtDescribe 데이터 품질이 표현 학습 및 일반화에 어떤 영향을 미치는가?

주요 결과

  • ProtST로 유도된 PLMs는 위치 특성화, 적합성 및 기능 주석 벤치마크에서 일관되게 일반 PLMs를 능가한다.
  • ProtST-ESM-2는 평가된 설정에서 가장 높은 localization 성능과 경쟁력 있는 기능 주석 지표를 달성한다.
  • 제로샷 ProtST 분류기는 위치화 및 반응 분류 과제에서 일부 소수샷 감독학습 벤치마크와 비슷하거나 이를 능가한다.
  • 제로샷 텍스트-단백질 검색은 기능 주석 없이 GO 프롬프트에서 기능적 단백질을 식별하는 능력을 보여준다.
  • 제로샷 및 감독 모델의 앙상블이 다운스트림 작업의 성능을 더 향상시킬 수 있다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.