Skip to main content
QUICK REVIEW

[论文解读] Learning Concave Conditional Likelihood Models for Improved Analysis of Tandem Mass Spectra

John T. Halloran, David M. Rocke|PubMed|Sep 4, 2019
Advanced Proteomics Techniques and Applications参考文献 24被引用 3
一句话总结

本文提出凸虚拟发射(CVEs),一种新型发射分布类别,可在动态贝叶斯网络中确保条件对数似然函数的凹性,从而在参数学习过程中实现全局收敛。通过将CVEs集成到Didea肽-谱匹配算法中,作者实现了最先进的评分准确率,并在推理速度上提升了64.2%,在1%假阳性率(FDR)下比DRIP和MS-GF+多识别出16%的谱图。

ABSTRACT

The most widely used technology to identify the proteins present in a complex biological sample is tandem mass spectrometry, which quickly produces a large collection of spectra representative of the <i>peptides</i> (i.e., protein subsequences) present in the original sample. In this work, we greatly expand the parameter learning capabilities of a dynamic Bayesian network (DBN) peptide-scoring algorithm, Didea [25], by deriving emission distributions for which its conditional log-likelihood scoring function remains concave. We show that this class of emission distributions, called <i>Convex Virtual Emissions</i> (CVEs), naturally generalizes the log-sum-exp function while rendering both maximum likelihood estimation and conditional maximum likelihood estimation concave for a wide range of Bayesian networks. Utilizing CVEs in Didea allows efficient learning of a large number of parameters while ensuring global convergence, in stark contrast to Didea's previous parameter learning framework (which could only learn a single parameter using a costly grid search) and other trainable models [12, 13, 14] (which only ensure convergence to local optima). The newly trained scoring function substantially outperforms the state-of-the-art in both scoring function accuracy and downstream Fisher kernel analysis. Furthermore, we significantly improve Didea's runtime performance through successive optimizations to its message passing schedule and derive explicit connections between Didea's new concave score and related MS/MS scoring functions.

研究动机与目标

  • 为克服以往Didea参数学习依赖昂贵网格搜索且仅能优化单一参数的局限性。
  • 开发一类通用的发射分布,确保在贝叶斯网络中实现凹条件对数似然,从而实现高效且全局收敛的参数学习。
  • 在不牺牲准确率的前提下,显著提升Didea的求和-乘积推理运行效率。
  • 通过推导条件对数似然梯度,增强下游判别性分析的性能,用于基于核函数的后处理方法。
  • 建立Didea新型评分函数与广泛使用的XCorr等方法之间的理论联系。

提出的方法

  • 通过求解一个非线性微分方程,推导出一类称为凸虚拟发射(CVEs)的发射分布,该方程强制一般贝叶斯网络发射分布满足凸性条件。
  • 证明CVEs推广了对数求和指数函数,并在最大似然与条件最大似然估计中均保持凹性。
  • 将CVEs集成到Didea动态贝叶斯网络模型中,实现对大量参数的高效、全局优化。
  • 通过算法优化改进Didea的消息传递调度,使运行时间减少64.2%。
  • 推导Didea评分的条件对数似然梯度,以丰富基于核函数的判别性后处理方法的特征空间。
  • 建立Didea评分与广泛使用的XCorr评分函数之间的理论下界,使未来可基于凹Didea框架对XCorr实现参数学习。

实验结果

研究问题

  • RQ1能否推导出一类通用的发射分布,确保在动态贝叶斯网络中实现凹条件对数似然,从而实现可扩展的参数学习?
  • RQ2如何将Didea的参数学习从单参数网格搜索扩展至支持多参数的高效、全局优化?
  • RQ3在不损害评分准确率的前提下,Didea的推理速度可提升至何种程度?
  • RQ4Didea条件对数似然的梯度是否能为后处理提供比现有方法更丰富的判别性特征?
  • RQ5Didea新型评分函数与已建立的MS/MS评分函数(如XCorr)之间是否存在理论关联?

主要发现

  • 所提出的凸虚拟发射(CVEs)类别推广了对数求和指数函数,并确保条件对数似然的凹性,从而在参数学习过程中实现全局收敛。
  • 采用CVEs的新Didea模型在基准数据集上,于严格的1% FDR条件下,比DRIP和MS-GF+多识别出16%的谱图。
  • 优化后的Didea实现使推理时间减少64.2%,每张谱图的搜索时间缩短至7秒以内,比DRIP快两个数量级。
  • Didea的条件对数似然梯度相比DRIP的对数似然梯度提供了更具信息量的特征,使Percolator能实现更优的校准性能。
  • 在1% FDR下,训练后的Didea评分函数比DRIP多识别出12.3%的谱图,比MS-GF+多识别出13.4%的谱图,在下游分析中全面超越所有现有最先进方法。
  • 建立了Didea评分与XCorr评分函数之间的理论下界,使未来可基于凹Didea框架对XCorr实现参数学习。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。