Skip to main content
QUICK REVIEW

[論文レビュー] DeepSF: deep convolutional neural network for mapping protein sequences to folds

Jie Hou, Badri Adhikari|arXiv (Cornell University)|Jun 4, 2017
Protein Structure and Dynamics参考文献 32被引用数 10
ひとこと要約

DeepSFは、従来のテンプレートベースの手法を回避して、タンパク質配列を1195の既知のフォールドの1つに直接マップする1次元畳み込みニューラルネットワークを導入する。SCOP1.75では80.4%の精度を達成し、SCOP2.06では77.0%を記録し、フォールド認識タスクにおいてHHSearchを4.5–29.1%上回り、配列の変異に対して耐性のある強固な特徴抽出を実現する。

ABSTRACT

Motivation Protein fold recognition is an important problem in structural bioinformatics. Almost all traditional fold recognition methods use sequence (homology) comparison to indirectly predict the fold of a tar get protein based on the fold of a template protein with known structure, which cannot explain the relationship between sequence and fold. Only a few methods had been developed to classify protein sequences into a small number of folds due to methodological limitations, which are not generally useful in practice. Results We develop a deep 1D-convolution neural network (DeepSF) to directly classify any protein se quence into one of 1195 known folds, which is useful for both fold recognition and the study of se quence-structure relationship. Different from traditional sequence alignment (comparison) based methods, our method automatically extracts fold-related features from a protein sequence of any length and map it to the fold space. We train and test our method on the datasets curated from SCOP1.75, yielding a classification accuracy of 80.4%. On the independent testing dataset curated from SCOP2.06, the classification accuracy is 77.0%. We compare our method with a top profile profile alignment method - HHSearch on hard template-based and template-free modeling targets of CASP9-12 in terms of fold recognition accuracy. The accuracy of our method is 14.5%-29.1% higher than HHSearch on template-free modeling targets and 4.5%-16.7% higher on hard template-based modeling targets for top 1, 5, and 10 predicted folds. The hidden features extracted from sequence by our method is robust against sequence mutation, insertion, deletion and truncation, and can be used for other protein pattern recognition problems such as protein clustering, comparison and ranking.

研究の動機と目的

  • シーケンスアラインメントやテンプレートマッチングに依存せずに、タンパク質配列を既知のフォールドに直接分類すること。
  • 相同性やテンプレート構造に依存する従来のフォールド認識手法の限界を克服すること。
  • 生タンパク質配列からフォールド関連特徴を直接学習できるディープラーニングモデルの開発。
  • 特にテンプレートフリーおよび困難なテンプレートベースのモデリング状況において、フォールド認識の精度を向上させること。
  • より広範なタンパク質パターン認識応用を想定した、変異に強く耐性のある特徴の生成。

提案手法

  • 任意の長さのタンパク質配列から階層的特徴を抽出できる1次元畳み込みニューラルネットワーク(CNN)を訓練する。
  • 学習された畳み込みフィルタを用いてアミノ酸配列を処理し、フォールドクラスを予測するのに有用な局所的およびグローバルなパターンを検出する。
  • 特徴はプーリングされ、全結合層を経由してSCOPが定義する1195のフォールドの1つを予測する。
  • エンドツーエンドでSCOP1.75のキュレート済みデータセットを用いて、交差エントロピー損失と確率的勾配降下法でネットワークを訓練する。
  • シーケンスの突然変異、挿入、欠失、切断に対する学習済み特徴の耐性を評価する。
  • CASP9–12のターゲットを用いて、HHSearchと比較してモデルの性能をベンチマークし、テンプレートベースおよびテンプレートフリーの予測タスクを含む。

実験結果

リサーチクエスチョン

  • RQ1ディープラーニングモデルは、シーケンスアラインメントや既知のテンプレートに依存せずに、タンパク質配列をフォールドに直接分類できるか?
  • RQ2提案手法の性能は、HHSearchのような最先端のプロファイルプロファイルアラインメント手法と比較して、フォールド認識においてどのように異なるか?
  • RQ3モデルが学習した特徴は、生物学的に重要なシーケンス変異に対してどれほど耐性があるか?
  • RQ4抽出された特徴は、クラスタリングやランク付けなどの他のタンパク質パターン認識タスクに一般化可能か?
  • RQ5新しいSCOPバージョンの独立したテストセットにおけるモデルの精度はいかほどか?

主な発見

  • DeepSFはSCOP1.75データセットで80.4%の分類精度を達成し、大規模なフォールド分類タスクにおいて優れた性能を示した。
  • SCOP2.06の独立したテストセットでも77.0%の高い精度を維持し、新しいデータへの一般化能力が優れていることが示された。
  • CASP9–12のテンプレートフリーモデリングターゲットにおいて、DeepSFはトップ1フォールド認識精度でHHSearchを14.5%~29.1%上回った。
  • 困難なテンプレートベースのモデリングターゲットでは、トップ1、トップ5、トップ10予測において、DeepSFはHHSearchよりも4.5%~16.7%の精度向上を達成した。
  • DeepSFが抽出する隠れ特徴は、シーケンスの突然変異、挿入、欠失、切断に対して耐性があり、下流タスクでの信頼できる利用が可能である。
  • モデルが学習した表現は、フォールド認識を超えて、タンパク質クラスタリング、比較、ランク付けの分野でも転送可能で効果的である。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。