Skip to main content
QUICK REVIEW

[Paper Review] Natively unfolded proteins: scalar predictors

Antonio Deiana, Andrea Giansanti|arXiv (Cornell University)|Jun 30, 2008
Protein Structure and Dynamics42 references1 citations
TL;DR

This study introduces and evaluates scalar predictors—mean packing (<P>), mean contact energy (<Ec>), and a novel gVSL2-based index—for identifying natively unfolded proteins. By combining these via strict unanimous (SSU) and comprehensive (S0) scoring schemes, the method achieves 79% sensitivity, 94% specificity, and only 6% false predictions, while revealing a scaling law with a critical exponent of 1.95 ± 0.21 in genome-wide unfolded protein frequency across domains of life.

ABSTRACT

This work revisits ab-initio methods to identify natively unfolded proteins. Single predictors and combined score indexes are considered and their performance is critically evaluated against other methods already present in the literature. We consider mean packing (< P>), mean contact energy(< Ec>) and a new index of folding status, based on VSL2 (gV SL2), a predictor of single disordered amino acids. We use a new dataset made of 743 folded proteins and 81 natively unfolded proteins. Individual use of these predictors has a performance comparable or even better than other proposed methods: gV SL2 reaches a sensitivity (Sn) of 0.81, a specificity (Sp) of 0.89 and a level of false predictions (fp) of 0.11. The performance of these single predictors is significantly improved if used in combination. We introduce a strictly unanimous combination score SSU and a new score S0, combining 10 dichotomic predictors. The former score leaves some sequences undecided, whereas the latter classifies with no exceptions all the sequences in a dataset. Through the combined use of both scores we get: Sn=0.79, Sp=0.94 and fp=0.06, with less than 6 % of proteins left unpredicted. The combined use of SSU and S0 applied to the problem of finding the frequency of occurrence of natively unfolded proteins in genomes from Nature’s three kingdoms gives the following figures: the percentage of natively unfolded proteins predicted by SSU are 4.1 % for Bacteria, 1.0 % for Archaea and 20.0 % for Eukarya; comparable, but not coincident with similar previous determinations. Evidence is given of a scaling law relating the number of natively unfolded proteins with the total number of proteins in a genome; a first estimate of the critical exponent is 1.95 ± 0.21. 2

Motivation & Objective

  • To improve ab-initio prediction of natively unfolded proteins using scalar predictors and combined scoring methods.
  • To evaluate the performance of single predictors like gVSL2, <P>, and <Ec> against existing methods.
  • To develop a combined prediction framework that minimizes false positives while maintaining high sensitivity.
  • To estimate the frequency of natively unfolded proteins across the three domains of life (Bacteria, Archaea, Eukarya).
  • To investigate a potential scaling law linking genome size to the number of natively unfolded proteins.

Proposed method

  • Uses a dataset of 743 folded and 81 natively unfolded proteins to train and test scalar predictors.
  • Employs gVSL2 as a novel index for folding status, based on single-amino-acid disorder prediction.
  • Introduces two combined scores: SSU (strictly unanimous) and S0 (comprehensive), integrating 10 dichotomic predictors.
  • Applies SSU and S0 in tandem to classify proteins, balancing prediction completeness and accuracy.
  • Analyzes genome-wide data from Nature’s three domains to estimate unfolded protein frequencies.
  • Fits a power-law model to determine the scaling exponent between total proteins and natively unfolded proteins.

Experimental results

Research questions

  • RQ1How do individual scalar predictors like <P>, <Ec>, and gVSL2 compare in performance to existing methods for identifying natively unfolded proteins?
  • RQ2Can combining multiple scalar predictors significantly improve prediction sensitivity and specificity?
  • RQ3What is the optimal balance between prediction completeness and accuracy when combining predictors?
  • RQ4What are the genome-wide frequencies of natively unfolded proteins in Bacteria, Archaea, and Eukarya?
  • RQ5Is there a scaling relationship between the total number of proteins and the number of natively unfolded proteins across genomes?

Key findings

  • The gVSL2 predictor alone achieves 81% sensitivity, 89% specificity, and only 11% false positive rate.
  • Combined use of SSU and S0 yields 79% sensitivity, 94% specificity, and just 6% false predictions, with less than 6% of proteins left undecided.
  • The SSU method predicts 4.1% of proteins as natively unfolded in Bacteria, 1.0% in Archaea, and 20.0% in Eukarya.
  • The observed frequencies are comparable but not identical to previous estimates, suggesting robustness and consistency.
  • A scaling law is identified with a critical exponent of 1.95 ± 0.21, indicating a power-law relationship between genome size and unfolded protein count.
  • The combined SSU and S0 framework enables reliable, scalable prediction across diverse genomes with minimal false positives.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.