[論文レビュー] Interpretable RNA Foundation Model from Unannotated Data for Highly Accurate RNA Structure and Function Predictions
RNA-FMは、自己教師あり学習を通じて23 millionの未注釈の非コードRNA配列で訓練されたファウンデーションモデルで、RNAの二次・3D構造と機能予測を改善する埋め込みを生成します。SARS-CoV-2解析を含み、タスク横断的な一般化能力が高い。
Non-coding RNA structure and function are essential to understanding various biological processes, such as cell signaling, gene expression, and post-transcriptional regulations. These are all among the core problems in the RNA field. With the rapid growth of sequencing technology, we have accumulated a massive amount of unannotated RNA sequences. On the other hand, expensive experimental observatory results in only limited numbers of annotated data and 3D structures. Hence, it is still challenging to design computational methods for predicting their structures and functions. The lack of annotated data and systematic study causes inferior performance. To resolve the issue, we propose a novel RNA foundation model (RNA-FM) to take advantage of all the 23 million non-coding RNA sequences through self-supervised learning. Within this approach, we discover that the pre-trained RNA-FM could infer sequential and evolutionary information of non-coding RNAs without using any labels. Furthermore, we demonstrate RNA-FM's effectiveness by applying it to the downstream secondary/3D structure prediction, SARS-CoV-2 genome structure and evolution prediction, protein-RNA binding preference modeling, and gene expression regulation modeling. The comprehensive experiments show that the proposed method improves the RNA structural and functional modelling results significantly and consistently. Despite only being trained with unlabelled data, RNA-FM can serve as the foundational model for the field.
研究の動機と目的
- ラベル依存モデルを超える構造と機能予測を改善するために、巨大な未注釈ncRNAデータの活用を動機づける。
- 23M ncRNA配列を用いた自己教師あり学習による、タスク非依存のRNAファウンデーションモデル(RNA-FM)の開発。
- RNA-FMの埋め込みが配列情報・構造情報・進化情報を捉えることを示す。
- 軽量なヘッドで微調整することで、複数の下流タスクで最先端の性能を達成できることを示す。
提案手法
- BERTアーキテクチャを基盤とした12層 Transformer(RNA-FM)を構築する。
- RNAcentralの23 million ncRNA配列を対象に、マスクトークン再構築(自己教師あり)でRNA-FMを事前学習する。
- 処理後、各配列をL×640の埋め込み行列として表現する。
- タスク固有のヘッドで微調整するか、下流モデルの特徴量として埋め込みを用いる。
- 多様なベンチマークで、主要な二次構造予測器および3D距離/閉包タスクと比較する。
- ウェブサーバを提供し、コード/ウェイトをコミュニティ利用向けに公開する。
実験結果
リサーチクエスチョン
- RQ1ラベルなしncRNAデータで学習したファウンデーションモデルが、構造的・機能的信号を捉える表現を学習できるか?
- RQ2RNA-FMの埋め込みは、二次構造予測・3D近接/距離予測・RNA-タンパク質相互作用・遺伝子調節タスクを、最新手法と比べて改善するか?
- RQ3RNA-FM表現には解釈可能な進化情報が埋め込まれているか?
- RQ4RNA-FMは規制領域やウイルスゲノム(例:SARS-CoV-2)へ、ベンチマーク全体でどれくらい一般化できるか?
主な発見
- RNA-FM埋め込みは、構造/機能特性に基づいて埋め込み空間内でncRNAタイプを整理し、学習された生物学的信号を示唆する。
- RNA-FMは二次構造ベンチマーク(例:ArchiveII600およびbpRNA TS0)で多くの最先端手法より高いF1を達成し、いくつかの設定でUFoldを上回ることがある。
- 3D近接において、RNA-FM埋め込みを用いたモデルはMSA共分散やPETfoldを用いたモデルを上回り、転移学習により小規模データセットで大きな改善を発揮する。
- RNA-FM埋め込みは、未注釈RNA配列のみで訓練されたにもかかわらず、タンパク質-RNA相互作用および遺伝子発現調節モデリングで競争力のある、または優れた性能を実現する。
- RNA-FMは3D距離予測タスクをサポートし、配列データと組み合わせるとR2とPMCCが向上しMSEが低下する。RNAパズルのエンドツーエンド微分可能な3D予測を可能にする。
- RNA-FM埋め込みはSARS-CoV-2ゲノム調節要素予測を強化し、ウイルス変異株間の進化トレンドを示すことができる。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。