Skip to main content
QUICK REVIEW

[论文解读] Multiple competition-based FDR control for peptide detection

Kristen Emery, Syamand Hasam|arXiv (Cornell University)|Jul 2, 2019
Advanced Proteomics Techniques and Applications参考文献 45被引用 5
一句话总结

本文提出了一种新颖的基于多重竞争的FDR控制框架,用于串联质谱中的肽段检测,通过为每个目标使用多个独立的假序列评分来提升统计效能。通过将单一假序列的目标-假序列竞争替换为跨多个假序列的统一竞争,该方法在低FDR阈值下可将肽段发现量提升高达50%,同时在有限样本中严格控制FDR。

ABSTRACT

Competition-based FDR control has been commonly used for over a decade in the computational mass spectrometry community (Elias and Gygi, 2007). Recently, the approach has gained significant popularity in other fields after Barber and Candes (2015) laid its theoretical foundation in a more general setting that included the feature selection problem. In both cases, the competition is based on a head-to-head comparison between an observed score and a corresponding decoy / knockoff. Keich and Noble (2017b) recently demonstrated some advantages of using multiple rather than a single decoy when addressing the problem of assigning peptide sequences to observed mass spectra. In this work, we consider a related problem -- detecting peptides based on a collection of mass spectra -- and we develop a new framework for competition-based FDR control using multiple null scores. Within this framework, we offer several methods, all of which are based on a novel procedure that rigorously controls the FDR in the finite sample setting. Using real data to study the peptide detection problem we show that, relative to existing single-decoy methods, our approach can increase the number of discovered peptides by up to 50% at small FDR thresholds.

研究动机与目标

  • 解决单一假序列目标-假序列竞争(TDC)在肽段检测中因高变异性与低统计效能而存在的局限性。
  • 开发一种新的FDR控制框架,利用每个目标假设的多个独立假序列评分以增强检测效能。
  • 在有限样本设定下,通过新颖的竞争为基础程序,严格控制假发现率。
  • 提出一种数据驱动方法(LBM),用于选择最优调参,以在保持FDR控制的前提下最大化发现量。
  • 在真实MS/MS数据上,证明该方法优于现有方法,包括单一假序列TDC和平均目标-假序列竞争(aTDC)。

提出的方法

  • 提出一种通用框架,基于每个假设使用多个零假设评分(假序列)进行竞争为基础的FDR控制,其中每个目标评分直接与多个假序列评分竞争。
  • 引入三种新方法——FDS、mirror和LBM,每种方法基于将多个假序列评分整合为单一竞争统计量的不同策略。
  • 在LBM中使用标记重抽样技术,通过测试直接最大化是否能控制FDR,来选择最优调参(c, λ),从而实现自适应性能。
  • 采用一种新颖的随机映射(mirandom),其效能优于均匀随机映射,且在模拟中被证明更优。
  • 应用一种新颖的有限样本FDR控制程序,假设假序列为i.i.d.或条件交换性,确保强FDR控制。
  • 使用模拟数据和真实MS/MS数据评估并比较所提方法与单一假序列TDC及aTDC的性能。

实验结果

研究问题

  • RQ1与单一假序列TDC相比,基于多重竞争的FDR控制是否能显著提升肽段发现数量?
  • RQ2当每个目标使用多个假序列评分时,所提框架是否能在有限样本中实现严格的FDR控制?
  • RQ3在不同实验条件下,所提方法中(FDS、mirror、LBM)哪一种在效能与FDR控制之间提供了最佳平衡?
  • RQ4在发现率与变异性方面,新方法与平均目标-假序列竞争(aTDC)相比表现如何?
  • RQ5LBM方法是否能以数据驱动方式有效选择最优调参,以在保持FDR控制的前提下最大化发现量?

主要发现

  • 所提出的基于多重竞争的FDR控制框架在真实MS/MS数据上,与单一假序列TDC相比,在低FDR阈值下可将肽段发现量提升高达50%。
  • LBM方法通过使用标记重抽样选择最优调参,在所有所提方法中提供了最佳的效能与FDR控制平衡。
  • 使用mirandom映射的mirror方法在效能上优于均匀随机映射,凸显了随机映射选择在竞争为基础FDR控制中的重要性。
  • FDS方法提供了具有强理论保证的确定性替代方案,尽管在实际应用中效能低于LBM。
  • 模拟与真实数据均表明,所提方法在有限样本中能严格控制FDR,即使在现实实验条件下亦成立。
  • 该框架具有广泛适用性,不仅限于肽段检测,还可用于基序发现与差异基因表达分析,尤其适用于排列或零假设评分有限的场景。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。