Skip to main content
QUICK REVIEW

[論文レビュー] Predicting Hydroxyl Mediated Nucleophilic Degradation and Molecular Stability of RNA Sequences through the Application of Deep Learning Methods

Ankit Singhal|arXiv (Cornell University)|Nov 9, 2020
RNA and protein synthesis mechanisms参考文献 22被引用数 6
ひとこと要約

本研究では、スタンフォードOpenVaccineデータセットを用いて、ヒドロキシル基を介した求核的分解とmRNA配列の化学的安定性を予測するための3つのディープラーニングモデル—LSTM、GRU、およびグラフ畳み込みネットワーク(GCN)—を提案および評価した。GRUモデルは、分解リスク予測において76%の精度と最小のRMSE(0.266)を達成し、安定なmRNA治療薬のインシリコスクリーニングの可能性を示した。

ABSTRACT

Synthesis and efficient implementation mRNA strands has been shown to have wide utility, especially recently in the development of COVID vaccines. However, the intrinsic chemical stability of mRNA poses a challenge due to the presence of 2'-hydroxyl groups in ribose sugars. The -OH group in the backbone structure enables a base-catalyzed nucleophilic attack by the deprotonated hydroxyl on the adjacent phosphorous and consequent self-hydrolysis of the phosphodiester bond. As expected for in-line hydrolytic cleavage reactions, the chemical stability of mRNA strands is highly dependent on external environmental factors, e.g. pH, temperature, oxidizers, etc. Predicting this chemical instability using a computational model will reduce the number of sequences synthesized and tested through identifying the most promising candidates, aiding the development of mRNA related therapies. This paper proposes and evaluates three deep learning models (Long Short Term Memory, Gated Recurrent Unit, and Graph Convolutional Networks) as methods to predict the reactivity and risk of degradation of mRNA sequences. The Stanford Open Vaccine dataset of 6034 mRNA sequences was used in this study. The training set consisted of 3029 of these sequences (length of 107 nucleotide bases) while the testing dataset consisted of 3005 sequences (length of 130 nucleotide bases), in structured (Lowest Entropy Base Pair Probability Matrix) and unstructured (Nodes and Edges) forms. The stability of mRNA strands was accurately generated, with the Graph Convolutional Network being the best predictor of reactivity ($RMSE = 0.249$) while the Gated Recurrent Unit Network was the best at predicting risks of degradation ($RMSE = 0.266$). Combining all target variables, the GRU performed the best with 76% accuracy. Results suggest these models can be applied to understand and predict the chemical stability of mRNA in the near future.

研究の動機と目的

  • 2'-ヒドロキシル基によるリン酸エステル結合の自己加水分解に起因する内在的mRNA不安定性の課題に対処すること。
  • mRNA治療薬開発における実験的負担を軽減し、配列安定性のインシリコ予測を可能にすること。
  • LSTM、GRU、GCNのディープラーニングモデルの性能を、変化する環境条件下での反応性および分解リスク予測において評価すること。
  • 多様なpH、温度、金属イオン条件下でのmRNA安定性を予測するための最も効果的なディープラーニングアーキテクチャを同定すること。
  • 合成前に不安定な配列をスクリーニングする際の誤差を最小限に抑えるために、バイナリ分類拡張を提案すること。

提案手法

  • スタンフォードOpenVaccineデータセットの6034個のmRNA配列を用いて、長短期記憶(LSTM)、ゲート付き再帰ユニット(GRU)、およびグラフ畳み込みネットワーク(GCN)モデルを訓練した。
  • 構造化データとして最低エントロピー塩基対確率行列を、GCNの入力としてノードとエッジとして非構造化データを用いた。
  • k分割交差検証を実施し、107ヌクレオチド(トレーニングセット)および130ヌクレオチド(テストセット)の配列でモデルを訓練した。
  • pH 10、pH 10にMg2+を添加、50°C、50°CにMg2+を添加するという複数の環境条件下で、RMSEおよびMAEを評価指標としてモデル最適化を行った。
  • 複数のターゲット変数の予測を統合し、全条件における平均RMSEおよびMAEを用いて全体的な性能を評価した。
  • 不安定な配列のスクリーニングにおける誤検出を最小限に抑えるために、バイナリ分類拡張を提案した。特に高リスク配列の同定に焦点を当てた。

実験結果

リサーチクエスチョン

  • RQ1ディープラーニングモデルは、変化する環境条件下でもmRNA配列の反応性および分解リスクを正確に予測できるか?
  • RQ2LSTM、GRU、GCNアーキテクチャは、mRNA安定性および分解傾向の予測においてどのように比較できるか?
  • RQ3複数の分解条件下で、どのディープラーニングモデルが最小の予測誤差(RMSEおよびMAE)を達成するか?
  • RQ4これらのモデルは、mRNA治療薬開発における実験的テストの必要性をどの程度低減できるか?
  • RQ5バイナリ分類拡張は、不安定なmRNA配列のスクリーニングの信頼性を向上させることができるか?

主な発見

  • グラフ畳み込みネットワーク(GCN)は、反応性予測において最高の性能を示し、RMSEが0.249であった。
  • ゲート付き再帰ユニット(GRU)モデルは、分解リスク予測において最良の性能を示し、RMSEが0.266であった。
  • 全ターゲット変数を統合した場合、GRUモデルは76%の最高全体精度を達成した。
  • GRUモデルは複数の条件下で最小のMAEを示し、大きな予測誤差が少ないことを示した。
  • GCNモデルは、複数の条件下でMAE値が低く、個々の大きな誤差に弱くないことが示された。
  • 強力な性能を示したが、予測値のスケール(0.5~4)のため、24%の平均絶対誤差率を示し、非自明な残差誤差が存在した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。