[논문 리뷰] Interpretable RNA Foundation Model from Unannotated Data for Highly Accurate RNA Structure and Function Predictions
RNA-FM은 자가지도 학습으로 주석되지 않은 2300만 개의 비암호화 RNA 시퀀스에서 학습된 기반 모델로, 임베딩을 생성하여 RNA 이차구조/3D 구조 및 기능 예측을 개선하며 SARS-CoV-2 분석 포함, 강력한 크로스태스크 일반화를 보임.
Non-coding RNA structure and function are essential to understanding various biological processes, such as cell signaling, gene expression, and post-transcriptional regulations. These are all among the core problems in the RNA field. With the rapid growth of sequencing technology, we have accumulated a massive amount of unannotated RNA sequences. On the other hand, expensive experimental observatory results in only limited numbers of annotated data and 3D structures. Hence, it is still challenging to design computational methods for predicting their structures and functions. The lack of annotated data and systematic study causes inferior performance. To resolve the issue, we propose a novel RNA foundation model (RNA-FM) to take advantage of all the 23 million non-coding RNA sequences through self-supervised learning. Within this approach, we discover that the pre-trained RNA-FM could infer sequential and evolutionary information of non-coding RNAs without using any labels. Furthermore, we demonstrate RNA-FM's effectiveness by applying it to the downstream secondary/3D structure prediction, SARS-CoV-2 genome structure and evolution prediction, protein-RNA binding preference modeling, and gene expression regulation modeling. The comprehensive experiments show that the proposed method improves the RNA structural and functional modelling results significantly and consistently. Despite only being trained with unlabelled data, RNA-FM can serve as the foundational model for the field.
연구 동기 및 목표
- 대용량의 주석되지 않은 ncRNA 데이터를 활용해 구조 및 기능 예측을 라벨 의존 모델 외로 개선하는 동기 부여.
- self-supervised learning on 23M ncRNA sequences로 task-agnostic RNA 기반 모델(RNA-FM) 개발.
- RNA-FM 임베딩이 순차적, 구조적, 진화 정보를 포착함을 입증.
- 경량 헤드를 사용한 미세조정으로 다수의 다운스트림 작업에서 최첨단 성능을 보여줌.
제안 방법
- BERT 아키텍처를 기반으로 한 12-layer 트랜스포머(RNA-FM) 구축.
- 23백만 ncRNA 시퀀스를 RNAcentral에서 마스킹 토큰 재구성(self-supervised)으로 사전 학습.
- 처리 후 각 시퀀스를 L x 640 임베딩 매트릭스로 표현.
- 작업별 헤드로 미세조정하거나 임베딩을 다운스트림 모델의 특징으로 사용.
- 다양한 벤치마크에서 선도하는 이차구조 예측기 및 3D 거리/폐쇄 작업과 비교.
- 웹 서버 제공 및 커뮤니티 사용을 위한 코드/가중치 공개.
실험 결과
연구 질문
- RQ1레이블이 없는 ncRNA 데이터로 학습된 기반 모델이 구조적 및 기능적 신호를 포착하는 표현을 학습할 수 있는가?
- RQ2RNA-FM 임베딩이 이차 구조 예측, 3D 근접/거리 예측, RNA-단백질/조절 작업을 최신 방법과 비교해 성능을 개선하는가?
- RQ3RNA-FM 표현에 해석가능한 진화 관련 정보가 내재되어 있는가?
- RQ4RNA-FM이 벤치마크 전반에서 규제 영역과 바이러스 게놈(예: SARS-CoV-2)에 얼마나 잘 일반화되는가?
주요 결과
- RNA-FM 임베딩은 구조/기능 특성에 따라 ncRNA 유형을 임베딩 공간에서 조직하여 학습된 생물학적 신호를 나타낸다.
- RNA-FM은 이차 구조 벤치마크에서 많은 최첨단 방법보다 높은 F1을 달성하고 UFold를 여러 설정에서 능가할 수 있다.
- 3D 근접의 경우 RNA-FM 임베딩을 사용하는 모델이 MSA 공변량 및 PETfold를 사용하는 모델보다 성능이 좋고, 전이 학습은 작은 데이터셋에서 큰 이득을 준다.
- RNA-FM 임베딩은 주석되지 않은 RNA 시퀀스만으로 학습되었음에도 불구하고 단백질-RNA 상호작용 및 유전자 발현 규제 모델링에서 경쟁력 있거나 우수한 성능을 발휘한다.
- RNA-FM 임베딩은 3D 거리 예측 작업을 지원하며 시퀀스 데이터와 결합하면 더 높은 R2 및 PMCC, 더 낮은 MSE를 달성하고 RNA 퍼즐에 대한 끝-대-끝 미분가능 3D 예측을 가능하게 할 수 있다.
- RNA-FM 임베딩은 SARS-CoV-2 게놈 규제 요소 예측을 향상시키고 바이러스 변이 간 진화 경향을 시사할 수 있다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.