[论文解读] A Unified Statistical Framework for RNA Sequence Data from Individual Cells and Tissue
该论文提出了一种统一的分层贝叶斯模型,联合分析单细胞和批量组织RNA-seq数据,通过EM与Gibbs采样结合的经验贝叶斯估计方法显式建模单细胞数据中的dropout事件。该框架在基因表达填补和细胞类型去卷积方面表现优异,在模拟实验和胎儿大脑数据应用中优于现有方法。
Recent advances in technology have enabled the measurement of RNA levels for individual cells. Compared to traditional tissue-level RNA-seq data, single cell sequencing yields valuable insights about gene expression profiles for different cell types, which is potentially critical for understanding many complex human diseases. However, developing quantitative tools for such data remains challenging because of high levels of technical noise, especially the events. A happens when the RNA for a gene fails to be amplified prior to sequencing, producing a false zero in the observed data. In this paper, we propose a unified statistical framework for both single cell and tissue RNA-seq data, formulated as a hierarchical model. Our framework borrows the strength from both data sources and carefully models the dropouts in single cell data, leading to a more accurate estimation of cell type specific gene expression profile. In addition, our model naturally provides inference on (i) the dropout entries in single cell data that need to be imputed for downstream analyses, and (ii) the mixing proportions of different cell types in tissue samples. We adopt an empirical Bayes approach, where parameters are estimated using the EM algorithm and approximate inference is obtained by Gibbs sampling. Simulation results illustrate that our framework outperforms existing approaches both in correcting for dropouts in single cell data, as well as in deconvolving tissue samples. We also demonstrate an application to gene expression data on fetal brains, where our model successfully imputes the dropout genes and reveals cell type specific expression patterns.
研究动机与目标
- 开发一种统一的统计框架,整合单细胞和批量组织RNA-seq数据,以改善基因表达分析。
- 解决单细胞RNA-seq中dropout事件带来的挑战,这些事件会产生虚假的零表达值,阻碍准确的表达谱分析。
- 通过跨数据类型借用信息,实现细胞类型特异性基因表达谱的精确估计。
- 为下游填补提供对单细胞数据中缺失(dropout)条目的合理推断。
- 利用单细胞参考谱系通过去卷积方法估计批量组织样本中的细胞类型组成。
提出的方法
- 构建一个分层贝叶斯模型,通过在单细胞和批量组织RNA-seq数据之间共享参数以提升估计精度。
- 在分层结构中显式使用零膨胀或dropout生成过程对单细胞数据中的dropout事件进行建模。
- 采用EM算法对超参数进行经验贝叶斯估计,利用来自两种数据源的观测数据。
- 使用Gibbs采样对潜在变量(包括dropout指示变量和细胞类型比例)进行后验推断。
- 将单细胞数据作为参考,用于去卷积混合组织样本并推断细胞类型组成。
- 利用两种数据类型的信息,增强对dropout事件的填补效果,提升表达谱估计的准确性。
实验结果
研究问题
- RQ1统一的统计模型能否有效结合单细胞和批量组织RNA-seq数据,以改善基因表达估计?
- RQ2如何利用批量组织数据的信息,准确建模并填补单细胞RNA-seq数据中的dropout事件?
- RQ3联合建模框架在批量组织样本中去卷积细胞类型比例方面的改进程度如何?
- RQ4所提出的框架是否在填补dropout事件和恢复真实基因表达谱方面优于现有方法?
- RQ5该模型能否在复杂组织(如胎儿大脑)中揭示具有生物意义的细胞类型特异性表达模式?
主要发现
- 在模拟研究中,所提出的框架在纠正单细胞RNA-seq数据中的dropout事件方面显著优于现有方法。
- 该模型能够对缺失(dropout)条目进行准确推断,为下游分析提供可靠的填补结果。
- 该框架成功实现了对批量组织样本中细胞类型比例的去卷积,其准确性优于基线方法。
- 在胎儿大脑数据的应用中,该模型通过有效填补dropout基因,成功识别出细胞类型特异性表达模式。
- 通过EM与Gibbs采样结合的经验贝叶斯方法,实现了对复杂高维数据的稳健且计算可行的推断。
- 对单细胞和批量数据的联合建模,相比单独分析任一数据类型,均能获得更精确的基因表达谱估计。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。