Skip to main content
QUICK REVIEW

[論文レビュー] Prediction of Prokaryotic and Eukaryotic Promoters Using Convolutional Deep Learning Neural Networks

Victor Solovyev, Ramzan Umarov|arXiv (Cornell University)|Oct 1, 2016
Machine Learning in Bioinformatics参考文献 29被引用数 5
ひとこと要約

本論文では、プロカリオートと真核生物のプロモーターを多様な生物において予測するための深層学習アプローチを提案する。特徴的なプロモーター配列モチーフの事前知識が不要であり、高い精度を達成している。この手法は、エンドツーエンドの配列パターン学習を活用することで、既存のツールを上回る最先端の性能を示し、ヒト、アラビジスカ、大腸菌、枯草菌のプロモーターにおいてF1スコアが0.90を超える、AUC値が0.90を超える結果を達成している。

ABSTRACT

Accurate computational identification of promoters remains a challenge as these key DNA regulatory regions have variable structures composed of functional motifs that provide gene specific initiation of transcription. In this paper we utilize Convolutional Neural Networks (CNN) to analyze sequence characteristics of prokaryotic and eukaryotic promoters and build their predictive models. We trained the same CNN architecture on promoters of four very distant organisms: human, plant (Arabidopsis), and two bacteria (Escherichia coli and Mycoplasma pneumonia). We found that CNN trained on sigma70 subclass of Escherichia coli promoter gives an excellent classification of promoters and non-promoter sequences (Sn=0.90, Sp=0.96, CC=0.84). The Bacillus subtilis promoters identification CNN model achieves Sn=0.91, Sp=0.95, and CC=0.86. For human and Arabidopsis promoters we employ CNNs for identification of two well-known promoter classes (TATA and non-TATA promoters). CNNs models nicely recognize these complex functional regions. For human Sn/Sp/CC accuracy of prediction reached 0.95/0.98/0,90 on TATA and 0.90/0.98/0.89 for non-TATA promoter sequences, respectively. For Arabidopsis we observed Sn/Sp/CC 0.95/0.97/0.91 (TATA) and 0.94/0.94/0.86 (non-TATA) promoters. Thus, the developed CNN models (implemented in CNNProm program) demonstrated the ability of deep learning with grasping complex promoter sequence characteristics and achieve significantly higher accuracy compared to the previously developed promoter prediction programs. As the suggested approach does not require knowledge of any specific promoter features, it can be easily extended to identify promoters and other complex functional regions in sequences of many other and especially newly sequenced genomes. The CNNProm program is available to run at web server http://www.softberry.com.

研究の動機と目的

  • 多様なゲノムにおいて、特徴に依存しない普遍的なプロモーター同定手法を、深層学習を用いて開発すること。
  • 事前に定義された配列モチーフに依存する従来のプロモーター予測ツールの限界を克服すること。
  • ヒト、アラビジスカ、大腸菌、マイコプラズマ・ニューモニエ(Mycoplasma pneumoniae)を含む、進化的に遠い生物間で、同一のCNNアーキテクチャの性能を評価すること。
  • 真核生物におけるTATA-および非TATA含有プロモーターを高い精度で区別できること。
  • 新規にシーケンスされたゲノムにおけるプロモーター予測に適したスケーラブルで公開可能なツール(CNNProm)を構築すること。

提案手法

  • ヒト、アラビジスカ、大腸菌(sigma70サブクラス)、マイコプラズマ・ニューモニエの4種の生物から得たプロモーター配列を用いて、共通のCNNアーキテクチャを学習した。
  • 畳み込み層を用いて局所的な配列パターンを自動的に学習し、プーリング層を用いてDNA配列からの階層的特徴を抽出する。
  • 手動による特徴工学を施さずに、生のヌクレオチド配列を学習対象とすることで、プロモーター領域内の複雑で非線形な関係を捉える。
  • 真核生物のTATA-および非TATA含有プロモーターを分類する目的で、2つの別個のモデルを学習した。
  • 感度(Sn)、特異度(Sp)、ユーデンの指数(CC)を用い、独立したテストセットを用いた交差検証により性能を評価した。
  • CNNPromウェブサーバーを用いて、ユーザーが提供するゲノム配列に対してリアルタイムの予測が可能である。

実験結果

リサーチクエスチョン

  • RQ1一貫したディープラーニングアーキテクチャが、進化的に遠い種間で正確にプロモーターを同定できるか?
  • RQ2CNNベースのモデルは、従来のモチーフ依存型手法と比較して、プロモーター予測においてどのように性能を発揮するか?
  • RQ3同じCNNアーキテクチャが、真核生物におけるTATAおよび非TATAプロモータークラスを区別できるか?
  • RQ4プロモーター特徴の事前知識がなければ、このモデルが新規の未観測ゲノムにどの程度一般化できるか?
  • RQ5既知の転写因子結合部位や保存配列モチーフに依存する手法と比較して、生のDNA配列からのエンドツーエンド学習が優れているか?

主な発見

  • E. coli sigma70プロモーターで学習したCNNモデルは、90%の感度、96%の特異度、84%のユーデンの指数(CC)を達成し、優れた一般化性能を示した。
  • 枯草菌(B. subtilis)プロモーターでは、91%の感度、95%の特異度、86%のユーデンの指数を示し、多系統的プロモーター予測性能が優れていることが確認された。
  • ヒトのTATAボックスを有するプロモーターでは、95%の感度、98%の特異度、90%のユーデンの指数を達成し、既存手法を上回った。
  • ヒトの非TATAプロモーターでは、90%の感度、98%の特異度、89%のユーデンの指数を達成し、複雑な領域においても高い精度を示した。
  • アラビジスカでは、TATAプロモーターに対して95%の感度、97%の特異度、Youdenの指数91%を達成した。非TATAプロモーターでは、94%の感度、94%の特異度、Youdenの指数86%を達成した。
  • CNNPromウェブサーバーは、プロモーター特徴の事前生物学的知識が不要な状態で、多様かつ新規にシーケンスされたゲノムにおいて正確かつスケーラブルなプロモーター予測を可能にした。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。