Skip to main content
QUICK REVIEW

[论文解读] SimCD: Simultaneous Clustering and Differential expression analysis for single-cell transcriptomic data

Seyednami Niyakan, Ehsan Hajiramezanali|arXiv (Cornell University)|Apr 4, 2021
Single-cell and spatial transcriptomics参考文献 38被引用 5
一句话总结

SimCD 是一种统一的贝叶斯方法,通过分层伽马-负二项分布(hGNB)模型,同时对单细胞RNA测序数据执行聚类和差异表达分析。通过在一个框架中建模细胞异质性和动态基因表达变化,SimCD 在检测已知和新型生物标志物方面优于独立的聚类与DE分析方法,且无需归一化或填补等预处理步骤。

ABSTRACT

Single-Cell RNA sequencing (scRNA-seq) measurements have facilitated genome-scale transcriptomic profiling of individual cells, with the hope of deconvolving cellular dynamic changes in corresponding cell sub-populations to better understand molecular mechanisms of different development processes. Several scRNA-seq analysis methods have been proposed to first identify cell sub-populations by clustering and then separately perform differential expression analysis to understand gene expression changes. Their corresponding statistical models and inference algorithms are often designed disjointly. We develop a new method -- SimCD -- that explicitly models cell heterogeneity and dynamic differential changes in one unified hierarchical gamma-negative binomial (hGNB) model, allowing simultaneous cell clustering and differential expression analysis for scRNA-seq data. Our method naturally defines cell heterogeneity by dynamic expression changes, which is expected to help achieve better performances on the two tasks compared to the existing methods that perform them separately. In addition, SimCD better models dropout (zero inflation) in scRNA-seq data by both cell- and gene-level factors and obviates the need for sophisticated pre-processing steps such as normalization, thanks to the direct modeling of scRNA-seq count data by the rigorous hGNB model with an efficient Gibbs sampling inference algorithm. Extensive comparisons with the state-of-the-art methods on both simulated and real-world scRNA-seq count data demonstrate the capability of SimCD to discover cell clusters and capture dynamic expression changes. Furthermore, SimCD helps identify several known genes affected by food deprivation in hypothalamic neuron cell subtypes as well as some new potential markers, suggesting the capability of SimCD for bio-marker discovery.

研究动机与目标

  • 解决现有单细胞RNA测序分析方法将聚类与差异表达(DE)分析作为独立、分离步骤所带来的局限性。
  • 开发一种统一的统计模型,联合推断细胞聚类和动态基因表达变化,同时考虑生物和实验噪声。
  • 通过整合基因水平和细胞水平因素,改进对零膨胀单细胞RNA测序数据的建模,无需归一化或填补。
  • 在异质性细胞群体中实现对条件特异性差异表达基因的稳健检测,特别是在食物剥夺等复杂条件下。
  • 通过整合分析促进在单细胞转录组数据中发现已知和新型生物标志物。

提出的方法

  • SimCD 采用分层伽马-负二项分布(hGNB)模型联合建模单细胞RNA测序计数数据,通过基因水平和细胞水平的随机效应捕捉过度离散和零膨胀特性。
  • 该模型在基因和细胞水平显式引入生物协变量,以控制处理条件和细胞类型等混杂因素,提升对技术变异的鲁棒性。
  • 采用高效的吉布斯采样推断算法,在单一统一的贝叶斯框架内估计潜在参数,包括细胞聚类分配和差异表达效应。
  • hGNB 模型通过建模不同条件(如对照组与食物剥夺组)下的基因表达变化,实现动态差异表达分析,同时识别细胞亚群。
  • SimCD 通过直接使用原始计数数据并结合统计严谨的分布框架,避免了归一化或填补等预处理步骤。
  • 该方法通过动态表达模式自然定义细胞异质性,从而获得更具生物学一致性的聚类和DE结果。

实验结果

研究问题

  • RQ1与顺序方法相比,统一的统计模型是否能同时提升单细胞RNA测序数据的聚类和差异表达分析性能?
  • RQ2在分层hGNB模型中同时建模基因水平和细胞水平因素,如何增强对生物相关差异表达和细胞聚类的检测能力?
  • RQ3在真实世界单细胞RNA测序数据集中,SimCD 在识别已知和新型生物标志物方面,相较于最先进方法的优越程度如何?
  • RQ4SimCD 是否能消除对归一化或填补等预处理步骤的需求,同时保持或提升分析性能?
  • RQ5SimCD 在捕获下丘脑神经元对食物剥夺等生物条件的动态表达变化方面表现如何?

主要发现

  • 在 CORTEX 数据集中,SimCD 的差异表达分析 AUC-ROC 为 0.3153,优于 ZINB-WaVE(0.2436)和 scVI(0.2977)。
  • 在 PBMC4k 数据集中,SimCD 的 AUC-ROC 为 0.4699,显著优于 ZINB-WaVE(0.4420)和 scVI(0.4163)。
  • 在下丘脑神经元数据集中,SimCD 检测到 10 个已知的食物剥夺响应基因,并识别出新型潜在标志物,其富集的顶级 GO 术语与囊泡和神经元细胞凋亡过程相关(如 GO:0097458,P = 5.9e-06)。
  • 与 DESingle 和 DESeq2 相比,SimCD 的 DE 结果更具生物学一致性,显著富集的术语与突触囊泡和细胞外细胞器相关(如 DESingle 的 GO:0031988,P = 1.3e-11)。
  • 消融研究显示,在 CORTEX 数据集上,不同潜在因子数(K)下,包含细胞水平协变量可提升聚类准确率,以调整轮廓系数(ASW)衡量。
  • SimCD 在检测 B 细胞(AUC = 0.7517 ± 0.0104)和树突状细胞(AUC = 0.8720 ± 0.0155)中的差异表达基因方面表现更优,优于 DESeq2 和 sigEMD。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。