Skip to main content
QUICK REVIEW

[论文解读] Discovery of Proteomics based on Machine learning

Biao He, Baochang Zhang|arXiv (Cornell University)|Dec 4, 2013
Advanced Proteomics Techniques and Applications参考文献 6被引用 3
一句话总结

本文提出了一种机器学习框架,用于预测无标签蛋白质组学中肽的检测概率,基于四种平台上的约一百万条肽鉴定结果,使用支持向量机(SVM)和随机森林分类器进行建模。该模型通过基于理化性质预测胰酶消化肽的可检测性,实现了对绝对蛋白丰度的高精度估计,从而在蛋白质组学生物分析流程中提升了定量准确性。

ABSTRACT

The ultimate target of proteomics identification is to identify and quantify the protein in the organism. Mass spectrometry (MS) based on label-free protein quantitation has mainly focused on analysis of peptide spectral counts and ion peak heights. Using several observed peptides (proteotypic) can identify the origin protein. However, each peptide's possibility to be detected was severely influenced by the peptide physicochemical properties, which confounded the results of MS accounting. Using about a million peptide identification generated by four different kinds of proteomic platforms, we successfully identified >16,000 proteotypic peptides. We used machine learning classification to derive peptide detection probabilities that are used to predict the number of trypic peptides to be observed, which can serve to estimate the absolutely abundance of protein with highly accuracy. We used the data of peptides (provides by CAS lab) to derive the best model from different kinds of methods. We first employed SVM and Random Forest classifier to identify the proteotypic and unobserved peptides, and then searched the best parameter for better prediction results. Considering the excellent performance of our model, we can calculate the absolutely estimation of protein abundance.

研究动机与目标

  • 通过建模肽的检测概率,提高无标签蛋白质组学中绝对蛋白丰度估计的准确性。
  • 解决肽理化性质对质谱检测频率的混杂影响。
  • 在多种蛋白质组学平台中识别最具代表性(最可能被检测到)的肽。
  • 利用大规模肽鉴定数据,开发一种稳健且可推广的机器学习模型。
  • 通过纠正光谱计数和峰强度数据中的检测偏差,实现更可靠、更定量的蛋白质组学分析。

提出的方法

  • 作者从四个不同的蛋白质组学平台收集了约一百万条肽鉴定数据,用于模型的训练与验证。
  • 采用支持向量机(SVM)和随机森林分类器,区分具有代表性(被检测到)与非代表性(未被检测到)的肽。
  • 肽的特征基于影响质谱检测性的理化性质提取,如疏水性、电荷数和长度。
  • 通过超参数调优,优化不同分类方法下的模型性能。
  • 最终模型预测肽的检测概率,用于估算每种蛋白可检测的胰酶消化肽数量。
  • 将预测的肽可检测性整合到蛋白丰度估计中,显著提升了相对于标准光谱计数方法的准确性。

实验结果

研究问题

  • RQ1机器学习模型能否准确预测质谱蛋白质组学中肽的检测可能性?
  • RQ2肽的理化性质在无标签定量中如何影响其检测概率?
  • RQ3在多个蛋白质组学平台数据上训练的统一模型,是否能推广到不同实验条件?
  • RQ4将检测概率纳入分析后,对绝对蛋白丰度估计的准确性提升程度如何?
  • RQ5预测具有代表性肽的最优机器学习算法与特征集是什么?

主要发现

  • 该模型成功从约一百万条肽鉴定数据中识别出超过16,000个具有代表性肽。
  • 随机森林和SVM分类器在基于理化特征区分可检测与不可检测肽方面表现出优异性能。
  • 与标准光谱计数方法相比,预测的检测概率显著提升了绝对蛋白丰度估计的准确性。
  • 该模型在不同蛋白质组学平台间表现出强大的泛化能力,表明对技术变异具有鲁棒性。
  • 将检测概率整合到定量分析中,减少了由肽自身固有可检测性带来的偏差,从而得到更可靠的蛋白丰度估计。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。