Skip to main content
QUICK REVIEW

[論文レビュー] 3D Protein Structure Predicted from Sequence

Debora S. Marks, Lucy J. Colwell|arXiv (Cornell University)|Oct 23, 2011
Protein Structure and Dynamics参考文献 67被引用数 4
ひとこと要約

本論文は、アミノ酸配列から3次元タンパク質構造を予測するデ・ノボ手法を提示する。複数配列アラインメントに適用されたデータ制約付き最大エントロピー・モデルを用い、進化的に共変する残基ペア(EICs)を同定し、それらを構造予測の長距離制約として用いる。この手法により、相同性モデリングや既知の構造テンプレートを用いずに、多様なタンパク質スーパーファミリーにおいてCα-RMSD誤差が2.7–5.1 Åの精度を達成した。

ABSTRACT

The evolutionary trajectory of a protein through sequence space is constrained by function and three-dimensional (3D) structure. Residues in spatial proximity tend to co-evolve, yet attempts to invert the evolutionary record to identify these constraints and use them to computationally fold proteins have so far been unsuccessful. Here, we show that co-variation of residue pairs, observed in a large protein family, provides sufficient information to determine 3D protein structure. Using a data-constrained maximum entropy model of the multiple sequence alignment, we identify pairs of statistically coupled residue positions which are expected to be close in the protein fold, termed contacts inferred from evolutionary information (EICs). To assess the amount of information about the protein fold contained in these coupled pairs, we evaluate the accuracy of predicted 3D structures for proteins of 50-260 residues, from 15 diverse protein families, including a G-protein coupled receptor. These structure predictions are de novo, i.e., they do not use homology modeling or sequence-similar fragments from known structures. The resulting low Cα-RMSD error range of 2.7-5.1Å, over at least 75% of the protein, indicates the potential for predicting essentially correct 3D structures for the thousands of protein families that have no known structure, provided they include a sufficiently large number of divergent sample sequences. With the current enormous growth in sequence information based on new sequencing technology, this opens the door to a comprehensive survey of protein 3D structures, including many not currently accessible to the experimental methods of structural genomics. This advance has potential applications in many biological contexts, such as synthetic biology, identification of functional sites in proteins and interpretation of the functional impact of genetic variants.

研究の動機と目的

  • アミノ酸配列データと進化的情報のみを用いて、デ・ノボで3次元タンパク質構造を予測する手法を開発すること。
  • 複数配列アラインメントにおける統計的共変性解析を通じて、折りたたまれたタンパク質内での空間的近接残基ペアを同定すること。
  • 共進化する残基ペアに含まれる進化的制約が、正確な3次元構造を再構築するのに十分であるかどうかを評価すること。
  • 特に相同性テンプレートを欠くタンパク質スーパーファミリーに対しても、構造予測を可能にすること。
  • 急速に拡大する配列データを活用して、ゲノム全体の構造的スクリーニングをスケーラブルに実施できるフレームワークを提供すること。

提案手法

  • 複数配列アラインメント内の観察された残基ペア頻度に制約を加えた最大エントロピー・モデルを構築する。
  • 逆イジングモデル(ポッツモデル)を用いてペアワイズ相互作用を推定し、統計的にカップリングされた残基ペア(進化的情報接触、EICs)を同定する。
  • 推定されたEICsを3次元構造予測パイプラインにおける長距離距離制約として適用する。
  • EIC制約と配列の好みに整合する3次元構造を生成するために、モンテカルロベースのサンプリング手法を用いる。
  • 既知のテンプレート、相同性モデリング、フラグメントライブラリに依存せずに、デ・ノボフォールディングを実行する。
  • 実験的に決定された構造との比較において、Cα-ルート・ミーン・スクエア・デバイエーション(RMSD)を用いて予測を検証する。

実験結果

リサーチクエスチョン

  • RQ1複数配列アラインメントにおける進化的共変性パターンが、デ・ノボで3次元タンパク質構造を再構築するのに十分な情報を提供できるか?
  • RQ2推定された残基接触(EICs)が、折りたたまれたタンパク質内での真の空間的近接性をどの程度反映しているか?
  • RQ3この手法は、既知の構造相同体を欠くタンパク質に対しても正確な3次元構造を予測できるか?
  • RQ4膜タンパク質(例:GPCR)を含む多様なタンパク質スーパーファミリーにおいて、構造予測の精度はどのように変動するか?
  • RQ5このアプローチを用いて信頼性のある3次元構造予測を達成するために、必要な最小の多様な配列数はどの程度か?

主な発見

  • 本手法は、タンパク質の骨格の少なくとも75%において、Cα-RMSD誤差が2.7 Åから5.1 Åの範囲で、デ・ノボで3次元タンパク質構造を予測することに成功した。
  • 本手法は、Gタンパク質共役受容体を含む15の多様なタンパク質スーパーファミリーに対しても正確な予測を達成し、挑戦的な膜タンパク質スーパーファミリーへの適用可能性を示した。
  • 予測された構造は、既知のテンプレートやフラグメントライブラリを一切使用せず、進化的共変性データのみに基づいている。
  • 本研究では、複数配列アラインメントに埋め込まれた進化的制約が、3次元フォールドトポロジーを推定するのに十分な情報を含んでいることを示した。
  • 十分な配列多様性があれば、本手法により、数千の未同定タンパク質スーパーファミリーの、大規模かつ高スループットな3次元構造予測が可能になることが示唆された。
  • 本手法は、50–260アミノ酸の幅広いタンパク質サイズおよびフォールドにわたり、強固であることが確認され、未同定ゲノム全体への広範な適用可能性を示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。