[论文解读] Fast reconstruction of compact context-specific metabolic networks via integration of microarray data
该论文提出FASTCORMICS,一种快速、鲁棒的工作流程,可将微阵列数据与基因组尺度代谢模型(GEMs)整合,以在几分钟内重构出紧凑、上下文特异的代谢网络。通过使用Barcode算法实现无阈值的基因表达离散化,并结合FASTCORE高效的线性规划方法,该方法生成了高质量的模型,在预测癌细胞系中必需基因方面表现出优于现有方法的准确性和鲁棒性。
Recently we proposed an algorithm for the fast reconstruction of compact context-specific metabolic networks (FASTCORE) that allowed dropping the reconstruction time to the time order of seconds (Vlassis et al.,2014). This extremely low computational demand opens new possibilities for improving the quality of the models. Several rounds of model reconstruction, testing of the model's predictions against real experimental data, curation steps of the input model and the set of core reactions as well as cross-validations assays are required to reconstruct high-quality models. These semi-automated model curations steps are in such extend not possible with competing algorithms due to their high computational demands. To adapt FASTCORE for the integration of microarray data, we therefore propose a new workflow: FASTCORMICS. FASTCORMICS requires as input microarray data and a Genome-scale reconstruction. FASTCORMICS is devoid of heuristic parameter settings and has a low computational demand with overall building times in the order of a few minutes. FASTCORMICS preprocesses the microarrays data with the discretization tool Barcode (Zillox et al, 2007). Barcode uses prior knowledge on the intensity distribution of each probe set for a given microarray platform to segregate between expressed genes and non-expressed genes. This preprocessing step allows circumventing the need of setting a heuristic expression threshold, which is critical for the output models as in response to this threshold alternative pathways or subsystems might be included or excluded, thereby heavily changing the functionalities of the model. In general, FASTCORMICS outperforms competing algorithms and allows obtaining high-quality, robust models in a high-throughput manner. This will allow the use of metabolic modelling as routine process for the analysis of microarray data e.g. in the field of personalized medicine.
研究动机与目标
- 解决从微阵列数据生成高质量、紧凑的上下文特异代谢模型的挑战。
- 克服现有方法中启发式表达阈值的局限性,这些阈值可能偏差模型结构和预测性能。
- 将模型重构的计算时间减少至数分钟,以支持迭代优化和高通量分析。
- 与通用GEMs及竞争算法相比,提升在癌细胞模型中必需基因识别的预测准确性。
- 开发一种鲁棒、无参数的工作流程,最大限度降低对微阵列数据中批次效应和平台特异性噪声的敏感性。
提出的方法
- 使用Barcode(一种基于探针强度分布先验知识的离散化工具)预处理微阵列数据,将基因分类为表达或不表达,避免使用任意的表达阈值。
- 通过基因-蛋白-反应(GPR)规则将Barcode识别的表达基因映射到反应,以定义模型重构的核心反应集。
- 应用FASTCORE(一种基于快速线性规划的算法)识别为实现所有核心反应通量所必需的最小非核心反应集合。
- 对非核心反应的包含施加惩罚,以确保模型紧凑、具有生物学相关性且复杂度最低。
- 可选地,通过介质组成和生物量功能对模型进行约束,以反映生理条件。
- 通过shRNA敲低数据进行必需性检测,并使用超几何检验评估DisGeNET中与肿瘤发生相关基因在预测必需基因中的富集程度。
实验结果
研究问题
- RQ1是否可以不依赖于任意表达阈值,将微阵列数据整合到基因组尺度代谢模型中?
- RQ2FASTCORMICS工作流程是否能生成比通用GEMs或竞争算法具有更高必需基因预测准确性的上下文特异模型?
- RQ3与基于阈值的方法相比,使用Barcode进行表达离散化在模型鲁棒性及对批次效应的敏感性方面有何影响?
- RQ4在计算效率和预测性能方面,FASTCORMICS相较于MBA、GIMME或IMAT等现有算法的优越程度如何?
- RQ5该工作流程是否可实现高通量应用,以支持个性化医疗等常规系统生物学应用?
主要发现
- FASTCORMICS在数分钟内即可重构上下文特异的代谢模型,与现有方法相比显著减少了计算时间。
- 针对cancer1(基于Recon 1)构建的模型在必需性测试中KS得分p值为8.1e-07,优于MBA算法的p值0.0044。
- cancer2模型(基于Recon 2)的KS得分p值为2.9e-04,表明其在必需基因预测方面具有强统计显著性。
- 超几何检验证实,所有模型中预测必需基因均显著富集了DisGeNET中的肿瘤发生相关基因,验证了其生物学相关性。
- cancer1模型包含188个核心反应和90个非核心反应,总计1168个反应和377个通量承载反应,体现了模型的紧凑性与功能一致性。
- 该工作流程依赖Barcode,消除了对启发式阈值的依赖,降低了模型偏差,提升了对数据变异的鲁棒性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。