Skip to main content
QUICK REVIEW

[論文レビュー] Natively unfolded proteins: scalar predictors

Antonio Deiana, Andrea Giansanti|arXiv (Cornell University)|Jun 30, 2008
Protein Structure and DynamicsBiochemistry, Genetics and Molecular Biology参考文献 42被引用数 1
ひとこと要約

本研究では、スカラー予測子(平均パッケージング<P>、平均接触エネルギー<Ec>、および新規のgVSL2ベースのインデックス)を導入し、ネイティブに折りたたまれないタンパク質を同定する手法を提案する。これらの予測子を厳密な一致(SSU)および包括的(S0)スコアリング方式で統合することで、感度79%、特異度94%、誤予測率6%を達成し、系統的分類の全ゲノムスケールで、臨界指数1.95 ± 0.21のスケーリング則が示された。

ABSTRACT

This work revisits ab-initio methods to identify natively unfolded proteins. Single predictors and combined score indexes are considered and their performance is critically evaluated against other methods already present in the literature. We consider mean packing (< P>), mean contact energy(< Ec>) and a new index of folding status, based on VSL2 (gV SL2), a predictor of single disordered amino acids. We use a new dataset made of 743 folded proteins and 81 natively unfolded proteins. Individual use of these predictors has a performance comparable or even better than other proposed methods: gV SL2 reaches a sensitivity (Sn) of 0.81, a specificity (Sp) of 0.89 and a level of false predictions (fp) of 0.11. The performance of these single predictors is significantly improved if used in combination. We introduce a strictly unanimous combination score SSU and a new score S0, combining 10 dichotomic predictors. The former score leaves some sequences undecided, whereas the latter classifies with no exceptions all the sequences in a dataset. Through the combined use of both scores we get: Sn=0.79, Sp=0.94 and fp=0.06, with less than 6 % of proteins left unpredicted. The combined use of SSU and S0 applied to the problem of finding the frequency of occurrence of natively unfolded proteins in genomes from Nature’s three kingdoms gives the following figures: the percentage of natively unfolded proteins predicted by SSU are 4.1 % for Bacteria, 1.0 % for Archaea and 20.0 % for Eukarya; comparable, but not coincident with similar previous determinations. Evidence is given of a scaling law relating the number of natively unfolded proteins with the total number of proteins in a genome; a first estimate of the critical exponent is 1.95 ± 0.21. 2

研究の動機と目的

  • スカラー予測子と統合スコアリング手法を用いて、ネイティブに折りたたまれないタンパク質のab-initio予測を改善すること。
  • gVSL2、<P>、<Ec>といった単一の予測子が、既存の手法と比較してどの程度の性能を示すかを評価すること。
  • 高い感度を維持しつつ、誤予測を最小限に抑える統合予測フレームワークを構築すること。
  • 細菌、古細菌、真核生物の3系統の全ゲノムスケールで、ネイティブに折りたたまれないタンパク質の頻度を推定すること。
  • ゲノムサイズとネイティブに折りたたまれないタンパク質数との間に、スケーリング関係が存在するかを調査すること。

提案手法

  • 743個の折りたたまれたタンパク質と81個のネイティブに折りたたまれないタンパク質から成るデータセットを用いて、スカラー予測子の学習とテストを行う。
  • gVSL2を、アミノ酸1残基単位での不規則性予測に基づく、折りたたみ状態の新しいインデックスとして採用する。
  • 2つの統合スコア、SSU(厳密な一致)とS0(包括的)を導入し、10個の二値予測子を統合する。
  • SSUとS0を併用してタンパク質を分類し、予測の包括性と正確性のバランスをとる。
  • Natureの3系統の全ゲノムデータを分析し、ネイティブに折りたたまれないタンパク質の頻度を推定する。
  • 全タンパク質数とネイティブに折りたたまれないタンパク質数との間のべき乗則モデルをフィットさせ、スケーリング指数を特定する。

実験結果

リサーチクエスチョン

  • RQ1gVSL2、<P>、<Ec>といった個々のスカラー予測子は、ネイティブに折りたたまれないタンパク質を同定する既存手法と比較して、どの程度の性能を示すか?
  • RQ2複数のスカラー予測子を統合することで、予測の感度と特異度が顕著に向上するか?
  • RQ3予測の包括性と正確性の最適なバランスをとるには、予測子をどのように統合すべきか?
  • RQ4細菌、古細菌、真核生物におけるネイティブに折りたたまれないタンパク質の全ゲノムスケールの頻度はどの程度か?
  • RQ5ゲノム全体において、全タンパク質数とネイティブに折りたたまれないタンパク質数との間に、スケーリング関係が存在するか?

主な発見

  • gVSL2予測子単体でも、感度81%、特異度89%、誤予測率11%を達成する。
  • SSUとS0の併用により、感度79%、特異度94%、誤予測率6%にまで向上し、未決定のタンパク質は6%未満にまで減少する。
  • SSU法では、細菌で4.1%、古細菌で1.0%、真核生物で20.0%のタンパク質がネイティブに折りたたまれないものと予測される。
  • 観察された頻度は、過去の推定値と比較して類似しているが同一ではないため、妥当性と一貫性が示された。
  • 臨界指数1.95 ± 0.21を有するスケーリング則が特定され、ゲノムサイズとネイティブに折りたたまれないタンパク質数との間にはべき乗則的関係が存在することが示された。
  • SSUとS0の統合フレームワークにより、誤予測を最小限に抑えつつ、多様なゲノムで信頼性の高いスケーラブルな予測が可能になった。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。