[论文解读] Statistical inference in partially observed branching processes with application to cell lineage tracking of in vivo hematopoiesis
本文提出一种基于矩量的M-估计量,用于部分观测的连续时间多类型分支过程,以实现对体内造血系统细胞谱系追踪数据的统计推断。通过推导高阶矩的闭式表达式,该方法能够高效估计复杂模型的参数——包括命运决定概率和自我更新率——并支持模型选择与交叉验证,应用于食蟹猴条形码数据揭示了比以往模型更细致的分化动力学。
Single-cell lineage tracking strategies enabled by recent experimental technologies have produced significant insights into cell fate decisions, but lack the quantitative framework necessary for rigorous statistical analysis of mechanistic models describing cell division and differentiation. In this paper, we develop such a framework with corresponding moment-based parameter estimation techniques for continuous-time, multi-type branching processes. Such processes provide a probabilistic model of how cells divide and differentiate, and we apply our method to study hematopoiesis, the mechanism of blood cell production. We derive closed-form expressions for higher moments in a general class of such models. These analytical results allow us to efficiently estimate parameters of much richer statistical models of hematopoiesis than those used in previous statistical studies. After validating the methodology in simulation studies, we apply our estimator to hematopoietic lineage tracking data from rhesus macaques. Our analysis provides a more complete understanding of cell fate decisions during hematopoiesis in non-human primates, which may be more relevant to human biology and clinical strategies than previous findings from murine studies. For example, in addition to previously estimated hematopoietic stem cell self-renewal rate, we are able to estimate fate decision probabilities and to compare structurally distinct models of hematopoiesis using cross validation. These estimates of fate decision probabilities and our model selection results should help biologists compare competing hypotheses about how progenitor cells differentiate. The methodology is transferrable to a large class of stochastic compartmental models and multi-type branching models, commonly used in studies of cancer progression, epidemiology, and many other fields.
研究动机与目标
- 解决在细胞条形码等实验系统中分析高分辨率单细胞谱系追踪数据时缺乏严谨统计框架的问题。
- 通过开发一种可计算的、基于矩量的推断方法,克服部分观测随机过程参数估计的计算不可行性。
- 在造血系统中超越简单的两 compartment 模型,实现对复杂参数(如命运决定概率和自我更新率)的估计。
- 提供一种可推广的统计框架,适用于系统生物学、肿瘤学和流行病学中的多类型分支过程。
- 通过交叉验证和基于矩量的准则,实现对结构各异的造血模型之间的模型选择。
提出的方法
- 提出一种基于相关性的M-估计量,其基础为广义矩量法(GMM),适用于连续时间、多类型的分支过程。
- 推导该过程高阶矩(如二阶与三阶矩)的闭式解析表达式,以实现高效估计。
- 以慢病毒条形码标记和高通量测序生成的谱系追踪细胞时间序列数据作为矩量估计的输入。
- 基于分支过程的理论矩建立矩条件,并求解估计方程以推断转移速率和命运概率。
- 应用交叉验证比较并选择竞争性的造血结构模型,平衡模型复杂度与拟合优度。
- 通过在各种模型误设和数据稀疏条件下进行广泛模拟研究,验证该方法。
实验结果
研究问题
- RQ1基于矩量估计是否可用于从时间序列谱系数据中推断部分观测的多类型分支过程中转移速率和命运决定概率?
- RQ2不同结构的造血模型(如具有不同数量祖细胞阶段的模型)在解释观测到的谱系动力学方面表现如何?
- RQ3所提出的方法在非人灵长类造血系统中对自我更新率和分化概率的估计能力如何?
- RQ4该估计量在面对模型误设、缺失数据和真实条形码实验中的实验噪声时,其稳健性如何?
- RQ5该方法能否支持模型选择,以区分造血系统中层级式与非层级式分化路径?
主要发现
- 该方法成功利用体内食蟹猴谱系追踪数据,估计了复杂多阶段造血分支模型中的自我更新率、命运决定概率和转移速率。
- 交叉验证表明,包含多个中间祖细胞类型的模型比更简单的两 compartment 模型更好地拟合数据,支持更复杂的分化架构。
- 推导出的闭式矩表达式使参数估计高效,无需依赖计算量大的基于模拟的方法(如 ABC)。
- 由于高阶矩的解析可计算性,该框架可估计以往难以获取的参数(如谱系特异性分化概率)。
- 模型选择结果与新兴的生物学证据一致,挑战了经典的造血层级模型,如早期祖细胞的多潜能性及谱系潜能的可塑性。
- 模拟研究在部分观测和测量误差等现实条件下表明,该方法对数据稀疏性和实验噪声具有稳健性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。