[论文解读] Causal network inference using biochemical kinetics
该论文提出了一种全贝叶斯框架——化学模型平均(CheMA),通过基于化学反应动力学推导出的非线性常微分方程(ODEs)来推断因果生化网络并预测动态行为。通过将反应图和动力学参数视为未知量并在其上进行平均,CheMA 实现了在真实反应网络未知情况下的稳健网络推断与动态预测,相较于线性模型,在模拟数据和真实磷酸化蛋白质组数据上均表现出更高的准确性。
Network models are widely used as structural summaries of biochemical systems. Statistical estimation of networks is usually based on linear or discrete models. However, the dynamics of these systems are generally nonlinear, suggesting that suitable nonlinear formulations may offer gains with respect to network inference and associated prediction problems. We present a general framework for both network inference and dynamical prediction that is rooted in nonlinear biochemical kinetics. This is done by considering a dynamical system based on a chemical reaction graph and associated kinetics parameters. Inference regarding both parameters and the reaction graph itself is carried out within a fully Bayesian framework. Prediction of dynamical behavior is achieved by averaging over both parameters and reaction graphs, allowing prediction even when the underlying reactions themselves are unknown or uncertain. Results, based on (i) data simulated from a mechanistic model of mitogen-activated protein kinase signaling and (ii) phosphoproteomic data from cancer cell lines, demonstrate that nonlinear formulations can yield gains in network inference and permit dynamical prediction in the challenging setting where the reaction graph is unknown.
研究动机与目标
- 开发一种基于非线性生化动力学的通用框架,用于联合进行网络推断与动态预测。
- 解决在底层反应图未知或不确定的情况下进行网络推断的挑战。
- 通过利用非线性 ODE 的结构非对称性,超越线性或离散模型,以推断因果关系。
- 在无需预设反应网络的前提下,实现对干预(如药物处理)下信号传导动态的预测建模。
- 提供一种贝叶斯方法,通过对反应图和动力学参数的不确定性进行积分,实现稳健的推断。
提出的方法
- 该方法使用基于化学反应图 $G$ 和动力学参数 $\bm{\theta}$ 推导出的 ODE 系统来建模生化动态,形式为 $d\bm{X}/dt = \bm{f}_G(\bm{X}, \bm{\theta})$。
- 采用全贝叶斯框架,将反应图 $G$ 和参数 $\bm{\theta}$ 均视为具有先验分布的隐变量。
- 通过在 $\bm{\theta}$ 上积分计算边缘似然 $p(\mathcal{D}|G)$,实现基于后验 $p(G|\mathcal{D}) \propto p(G) \cdot p(\mathcal{D}|G)$ 的模型比较与选择。
- CheMA 1.0 通过采用米氏-门特恩动力学和吉布斯内梅特罗波利斯采样,高效探索 $G$ 和 $\bm{\theta}$ 的联合空间。
- 通过在多个反应图和参数集合上进行平均,获得动态预测,从而在模型不确定性下提升鲁棒性。
- 该方法利用时间序列数据,同时推断网络结构(通过粗粒度网络 $N$ 中边的后验概率边际)和系统动态。
实验结果
研究问题
- RQ1基于非线性 ODE 的全贝叶斯框架,相较于线性或离散模型,是否能提升因果网络推断的性能?
- RQ2当真实反应图未知或不确定时,能在多大程度上实现动态预测?
- RQ3为何利用化学动力学能够识别因果关系,即使非线性模型具有结构非对称性?
- RQ4当动力学参数从数据中难以识别时,CheMA 1.0 是否仍能准确推断网络结构?
- RQ5在真实磷酸化蛋白质组数据和模拟信号传导系统上,CheMA 的性能与现有方法相比如何?
主要发现
- 尽管单个动力学参数的可识别性较差,CheMA 1.0 在基于丝裂原活化蛋白激酶信号传导模型的模拟数据中仍成功识别出正确的因果网络结构。
- 在癌症细胞系的真实磷酸化蛋白质组数据上,CheMA 在网络准确性与动态预测方面均优于现有的线性与离散网络推断方法。
- 该方法生成了物理上合理的动态轨迹——平滑、有界且非负,表明动力学模型预测结果具有可解释性。
- 即使动力学参数无法可靠估计,因果网络结构仍可被识别,表明网络推断是完整动力系统的一个低维、更鲁棒的投影。
- CheMA 可在缺乏已知反应图的情况下,预测干预(如药物处理)下的动态行为,展示了其在不确定性下的系统生物学应用潜力。
- 计算成本较高,27种蛋白质系统需超过12小时,限制了其在极高维网络中的可扩展性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。