Skip to main content
QUICK REVIEW

[论文解读] Automated adaptive inference of coarse-grained dynamical models in systems biology

Bryan C. Daniels, Ilya Nemenman|arXiv (Cornell University)|Apr 24, 2014
Gene Regulatory Network Analysis参考文献 39被引用 6
一句话总结

本文提出了一种自适应、计算高效的系统生物学粗粒度动力学模型推断方法,通过自动调整模型复杂度以匹配数据可用性。利用分层嵌套的模型空间和基于马尔可夫链蒙特卡洛的贝叶斯模型选择,该方法可避免过拟合,并在数据有限且存在未观测变量的情况下实现准确预测,如在行星运动、信号转导通路和酿酒酵母糖酵解模型中的验证所示。

ABSTRACT

Cellular regulatory dynamics is driven by large and intricate networks of interactions at the molecular scale, whose sheer size obfuscates understanding. In light of limited experimental data, many parameters of such dynamics are unknown, and thus models built on the detailed, mechanistic viewpoint overfit and are not predictive. At the other extreme, simple ad hoc models of complex processes often miss defining features of the underlying systems. Here we propose an approach that instead constructs phenomenological, coarse-grained models of network dynamics that automatically adapt their complexity to the amount of available data. Such adaptive models lead to accurate predictions even when microscopic details of the studied systems are unknown due to insufficient data. The approach is computationally tractable, even for a relatively large number of dynamical variables, allowing its software realization, named Sir Isaac, to make successful predictions even when important dynamic variables are unobserved. For example, it matches the known phase space structure for simulated planetary motion data, avoids overfitting in a complex biological signaling system, and produces accurate predictions for a yeast glycolysis model with only tens of data points and over half of the interacting species unobserved.

研究动机与目标

  • 解决由于系统生物学中实验数据不足导致详细机理模型过拟合的问题。
  • 开发一种可根据数据可用性自动选择模型复杂度的方法,避免依赖微观细节。
  • 提出一种计算上可行的方法,从时间序列数据中推断动力学系统,即使关键变量未被观测到。
  • 通过使用嵌套的、完整的现象学模型层次结构,确保统计一致性并避免过拟合。
  • 构建尽可能简单的可解释、可预测模型,但不过分简化至低于数据所要求的程度。

提出的方法

  • 该方法使用嵌套的、完整的粗粒度动力学模型层次结构,其中每个模型均为下一个模型的简化版本,从而确保统计一致性。
  • 通过贝叶斯模型证据进行模型选择,每个模型的对数似然通过拉普拉斯近似方法计算边际似然。
  • 采用混合蒙特卡洛方法(Metropolis-Hastings)从模型参数的后验分布中抽样,其提议分布基于卡方函数的海森矩阵自适应调整。
  • 使用Levenberg-Marquardt算法进行局部参数优化,以最小化归一化卡方值,通过梯度阈值检查收敛性。
  • 通过温度调节的似然函数平衡MCMC抽样中的探索与接受率,参数先验由初始点处的海森矩阵决定。
  • 模型评估包括按模型规模统计的累积积分次数,计算成本在模型复杂度饱和后呈线性增长。

实验结果

研究问题

  • RQ1是否存在一种模型推断方法,能够根据可用数据量自动调整其复杂度,从而在观测数据有限的系统中避免过拟合?
  • RQ2如何构建现象学模型,使其在关键变量未被观测的情况下仍保持可解释性,同时准确捕捉复杂动力学?
  • RQ3在具有非线性动力学的系统中,采用贝叶斯模型选择的分层嵌套模型空间是否能实现一致且具有预测能力的推断?
  • RQ4当数据稀疏或噪声较大时,该方法在多大程度上优于传统机理建模方法?
  • RQ5该方法是否能仅从时间序列数据中可靠地恢复已知的动力学结构(例如相空间),而无需预先了解底层机制?

主要发现

  • 该方法仅使用时间序列数据,成功恢复了模拟行星运动的已知相空间结构,证明了其在定性预测上的准确性。
  • 在复杂的生物信号转导系统中,自适应模型避免了过拟合,并在高维、噪声较大的数据下产生了可靠预测。
  • 对于仅含30–50个数据点且超过50%物种未被观测的酿酒酵母糖酵解模型,该方法通过选择约65个参数的模型实现了准确预测。
  • 当模型复杂度随数据量增加而提升时,模型评估次数呈超线性增长;但一旦模型规模饱和,计算成本呈线性增长,表明计算效率较高。
  • 基于海森矩阵特征值大于1的有效参数数量始终低于数据点数量,证实了模型的一致性并避免了过拟合。
  • 该方法已通过名为Sir Isaac的软件实现,在包括存在未观测变量和非线性动力学的多种系统中表现出稳健性能。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。