[论文解读] Natively unfolded proteins: scalar predictors
本研究提出并评估了标量预测器——平均堆积度(<P>)、平均接触能(<Ec>)以及一种基于gVSL2的新索引——以识别天然无折叠蛋白。通过结合这些预测器,采用严格一致(SSU)和综合(S0)评分方案,该方法实现了79%的敏感度、94%的特异度,且仅6%的假阳性预测率,同时揭示了在所有生命域中基因组范围的无折叠蛋白频率存在临界指数为1.95 ± 0.21的标度律。
This work revisits ab-initio methods to identify natively unfolded proteins. Single predictors and combined score indexes are considered and their performance is critically evaluated against other methods already present in the literature. We consider mean packing (&lt; P&gt;), mean contact energy(&lt; Ec&gt;) and a new index of folding status, based on VSL2 (gV SL2), a predictor of single disordered amino acids. We use a new dataset made of 743 folded proteins and 81 natively unfolded proteins. Individual use of these predictors has a performance comparable or even better than other proposed methods: gV SL2 reaches a sensitivity (Sn) of 0.81, a specificity (Sp) of 0.89 and a level of false predictions (fp) of 0.11. The performance of these single predictors is significantly improved if used in combination. We introduce a strictly unanimous combination score SSU and a new score S0, combining 10 dichotomic predictors. The former score leaves some sequences undecided, whereas the latter classifies with no exceptions all the sequences in a dataset. Through the combined use of both scores we get: Sn=0.79, Sp=0.94 and fp=0.06, with less than 6 % of proteins left unpredicted. The combined use of SSU and S0 applied to the problem of finding the frequency of occurrence of natively unfolded proteins in genomes from Nature’s three kingdoms gives the following figures: the percentage of natively unfolded proteins predicted by SSU are 4.1 % for Bacteria, 1.0 % for Archaea and 20.0 % for Eukarya; comparable, but not coincident with similar previous determinations. Evidence is given of a scaling law relating the number of natively unfolded proteins with the total number of proteins in a genome; a first estimate of the critical exponent is 1.95 ± 0.21. 2
研究动机与目标
- 通过标量预测器和组合评分方法,改进天然无折叠蛋白的从头预测。
- 评估单个预测器(如gVSL2、<P>和<Ec>)相较于现有方法在识别天然无折叠蛋白方面的性能。
- 开发一种组合预测框架,在保持高敏感度的同时最大限度减少假阳性。
- 估算三个生命域(细菌、古菌、真核生物)中天然无折叠蛋白的频率。
- 探究基因组大小与天然无折叠蛋白数量之间潜在的标度关系。
提出的方法
- 使用包含743个折叠蛋白和81个天然无折叠蛋白的数据集,训练和测试标量预测器。
- 采用gVSL2作为基于单氨基酸残基紊乱预测的新型折叠状态索引。
- 引入两种组合评分:SSU(严格一致)和S0(综合),整合10个二元预测器。
- 并行应用SSU和S0对蛋白质进行分类,以平衡预测的完整性与准确性。
- 分析来自《自然》杂志所涵盖的三个生命域的基因组范围数据,以估算无折叠蛋白的频率。
- 拟合幂律模型,以确定总蛋白数与天然无折叠蛋白数之间的标度指数。
实验结果
研究问题
- RQ1单个标量预测器(如<P>、<Ec>和gVSL2)在识别天然无折叠蛋白方面的性能与现有方法相比如何?
- RQ2结合多个标量预测器是否能显著提升预测的敏感度与特异度?
- RQ3在组合预测器时,如何实现预测完整性与准确性的最佳平衡?
- RQ4细菌、古菌和真核生物中天然无折叠蛋白的基因组范围频率是多少?
- RQ5在基因组范围内,总蛋白数与天然无折叠蛋白数之间是否存在标度关系?
主要发现
- 仅使用gVSL2预测器即可实现81%的敏感度、89%的特异度,且假阳性率仅为11%。
- SSU与S0联合使用可实现79%的敏感度、94%的特异度,假阳性预测率仅为6%,且不足6%的蛋白质处于未决定状态。
- SSU方法在细菌中预测4.1%的蛋白质为天然无折叠,在古菌中为1.0%,在真核生物中为20.0%。
- 观察到的频率与先前估计值相近但不完全相同,表明结果具有稳健性与一致性。
- 识别出一条标度律,其临界指数为1.95 ± 0.21,表明基因组大小与无折叠蛋白数量之间存在幂律关系。
- SSU与S0组合框架可在多种基因组中实现可靠且可扩展的预测,且假阳性极少。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。