Skip to main content
QUICK REVIEW

[论文解读] SL$^2$MF: Predicting Synthetic Lethality in Human Cancers via Logistic Matrix Factorization

Yong Liu, Min Wu|arXiv (Cornell University)|Oct 20, 2018
Bioinformatics and Genomic Networks被引用 10
一句话总结

本文提出SL$^2$MF,一种逻辑矩阵分解方法,通过从观测到的SL数据中学习潜在基因表征,结合重要性加权处理已知SL对,并整合蛋白质-蛋白质相互作用(PPI)网络和基因本体(GO)的生物学知识,以预测人类癌症中的合成致死性。该方法在两种验证场景中分别取得了0.7330和0.6701的AUC分数,表现出优异性能,可作为识别新型抗癌药物靶点的补充工具。

ABSTRACT

Synthetic lethality (SL) is a promising concept for novel discovery of anti-cancer drug targets. However, wet-lab experiments for detecting SLs are faced with various challenges, such as high cost, low consistency across platforms or cell lines. Therefore, computational prediction methods are needed to address these issues. This paper proposes a novel SL prediction method, named SL2MF, which employs logistic matrix factorization to learn latent representations of genes from the observed SL data. The probability that two genes are likely to form SL is modeled by the linear combination of gene latent vectors. As known SL pairs are more trustworthy than unknown pairs, we design importance weighting schemes to assign higher importance weights for known SL pairs and lower importance weights for unknown pairs in SL2MF. Moreover, we also incorporate biological knowledge about genes from protein-protein interaction (PPI) data and Gene Ontology (GO). In particular, we calculate the similarity between genes based on their GO annotations and topological properties in the PPI network. Extensive experiments on the SL interaction data from SynLethDB database have been conducted to demonstrate the effectiveness of SL2MF.

研究动机与目标

  • 为解决湿实验筛选人类癌症中合成致死(SL)互作的高成本和不一致性问题。
  • 通过开发一种计算方法,更有效地利用已知SL对,以克服可靠人类SL数据稀缺的问题。
  • 通过整合蛋白质-蛋白质相互作用(PPI)网络和基因本体(GO)注释的生物学知识,提升预测准确性。
  • 为人类癌症基因组学中的未来SL预测方法提供一个稳健、数据驱动的基线。

提出的方法

  • SL$^2$MF采用逻辑矩阵分解,将基因对之间合成致死性的概率建模为学习到的基因特异性潜在向量的线性组合。
  • 在训练过程中应用重要性加权,优先考虑已知SL对,为可靠互作分配更高权重,为未知对分配较低权重。
  • 通过基因本体语义相似性和PPI网络的拓扑特征计算基因相似性,以丰富潜在向量的学习。
  • 模型在SynLethDB数据库的SL数据上进行训练,包含两种评估场景:一种排除DAISY预测的SL对,另一种包含它们以评估泛化能力。
  • 最终预测按估计互作概率排序,性能通过AUC和AUPR指标评估。
  • 在相同设置下与最先进的SL预测方法DAISY进行比较,以验证其有效性。

实验结果

研究问题

  • RQ1逻辑矩阵分解能否有效利用人类癌症中不完整且嘈杂的SL数据来建模合成致死性概率?
  • RQ2与均匀加权相比,对已知SL对进行重要性加权在多大程度上提升了预测性能?
  • RQ3整合PPI网络拓扑结构和GO注释在多大程度上提升了SL预测的准确性?
  • RQ4SL$^2$MF在预测性能和前几名预测结果的重叠度方面,与现有方法如DAISY相比如何?
  • RQ5SL$^2$MF能否作为人类癌症基因组学中未来SL预测方法的可靠基线?

主要发现

  • 在场景1中,SL$^2$MF的AUC达到0.7330,该场景排除了DAISY预测的SL对,表明其具备强大的泛化能力。
  • 在场景2中,由于训练集中包含DAISY预测的SL对,数据多样性降低且可能存在过拟合,AUC下降至0.6701。
  • 两种场景的AUPR分数均低于0.005,表明许多高排名预测为假阳性,但模型仍捕捉到了部分真实SL对。
  • 在前5,740个预测的SL对中,场景1有14对在SynLethDB数据库中得到验证,而DAISY无任何重叠,表明SL$^2$MF在此设置下具有更优的召回能力。
  • 在场景2中,SL$^2$MF的前27个预测对经SynLethDB和文本挖掘验证,而DAISY仅验证了3对,凸显SL$^2$MF在预测可靠性方面的提升。
  • SL$^2$MF与DAISY预测结果的重叠极小——场景1中仅有5对共同预测,场景2中仅3对,表明两者具有不同的预测机制且具有互补性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。