[论文解读] Supervised Classification of Flow Cytometric Samples via the Joint Clustering and Matching (JCM) Procedure
本文提出了一种基于联合聚类与匹配(JCM)流程的监督分类方法,用于流式细胞术样本分类。该方法通过偏斜混合模型对细胞群进行建模,并通过最小化新样本拟合密度与类别模板之间的Kullback-Leibler散度实现分类。JCM在急性髓系白血病(AML)分类挑战中表现优异,AUC达到完美值,灵敏度为100%,优于五种基准方法。
We consider the use of the Joint Clustering and Matching (JCM) procedure for the supervised classification of a flow cytometric sample with respect to a number of predefined classes of such samples. The JCM procedure has been proposed as a method for the unsupervised classification of cells within a sample into a number of clusters and in the case of multiple samples, the matching of these clusters across the samples. The two tasks of clustering and matching of the clusters are performed simultaneously within the JCM framework. In this paper, we consider the case where there is a number of distinct classes of samples whose class of origin is known, and the problem is to classify a new sample of unknown class of origin to one of these predefined classes. For example, the different classes might correspond to the types of a particular disease or to the various health outcomes of a patient subsequent to a course of treatment. We show and demonstrate on some real datasets how the JCM procedure can be used to carry out this supervised classification task. A mixture distribution is used to model the distribution of the expressions of a fixed set of markers for each cell in a sample with the components in the mixture model corresponding to the various populations of cells in the composition of the sample. For each class of samples, a class template is formed by the adoption of random-effects terms to model the inter-sample variation within a class. The classification of a new unclassified sample is undertaken by assigning the unclassified sample to the class that minimizes the Kullback-Leibler distance between its fitted mixture density and each class density provided by the class templates.
研究动机与目标
- 为解决将新流式细胞术样本分类到预定义疾病或健康结局类别的挑战。
- 克服手动门控和基于特征的分类器在高维、复杂流式细胞术数据中的局限性。
- 开发一种基于参数模型的分类方法,利用完整密度信息而非摘要特征。
- 在真实世界数据集上,评估基于JCM的分类器与现有最先进方法的性能表现。
提出的方法
- 使用有限混合的偏斜分布(例如受限偏斜正态分布或t分布)对每个样本中的细胞群进行建模,以捕捉偏态和重尾特征。
- 通过JCM框架同时执行聚类与样本间聚类匹配,实现多个样本间细胞群的对齐。
- 通过在分层混合模型中引入随机效应项,对样本间变异进行建模,构建每个预定义类别的类别模板。
- 通过计算新未分类样本的拟合混合密度与每个类别模板密度之间的Kullback-Leibler(KL)散度,实现分类。
- 使用期望最大化(EM)算法估计模型参数,包括分量均值、协方差、混合比例及随机效应项。
- 在完整模型拟合前,应用降维(主成分分析、非负矩阵分解或广义矩阵分解)或异常值剔除(通过在标记子集上应用JCM)以提升鲁棒性与可扩展性。
实验结果
研究问题
- RQ1JCM流程能否有效适应于将流式细胞术样本分类到预定义类别的监督分类任务?
- RQ2基于JCM的分类性能与现有基于特征的分类器(如SVM、Citrus和HDPGMM)相比如何?
- RQ3通过KL散度进行完整密度比较是否优于仅依赖摘要统计量(如聚类比例)的方法?
- RQ4JCM方法在高维数据及临床流式细胞术数据集中样本间变异情况下的鲁棒性如何?
主要发现
- 在AML分类挑战中,JCM的受试者工作特征曲线下面积(AUC)接近1.0,为六种比较方法中的最高值。
- JCM正确分类了所有AML样本,灵敏度达到1.0,而其他任何方法的灵敏度均未超过0.95。
- 在BCR数据集中,JCM在F-measure和AUC两项指标上均优于所有其他方法,展现出更高的分类准确性。
- 采用完整参数化密度模型与KL散度进行比较,相比依赖有限统计量的基于特征的方法,实现了更精确的分类。
- 在类别模板中集成随机效应项有效捕捉了样本间的变异,提升了同一类别内样本间的泛化能力。
- 经过异常值剔除和降维等预处理步骤后,该方法在高维设置下依然保持鲁棒性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。