Skip to main content
QUICK REVIEW

[論文レビュー] Supervised Classification of Flow Cytometric Samples via the Joint Clustering and Matching (JCM) Procedure

Sharon Lee, Geoffrey J. McLachlan|arXiv (Cornell University)|Nov 11, 2014
Single-cell and spatial transcriptomics参考文献 24被引用数 6
ひとこと要約

本論文は、歪み混合モデルを用いて細胞集団をモデル化し、新しいサンプルを最小化するKullback-Leibler発散に基づいて分類する、連合クラスタリングおよびマッチング(JCM)手順を用いた、フローサイトメトリー・サンプルの教師あり分類手法を提案する。JCMは優れた性能を示し、急性前骨髄性白血病(AML)分類チャレンジにおいて、AUCが1.0に近く、100%の感度を達成し、比較された5つのベンチマーク手法を上回った。

ABSTRACT

We consider the use of the Joint Clustering and Matching (JCM) procedure for the supervised classification of a flow cytometric sample with respect to a number of predefined classes of such samples. The JCM procedure has been proposed as a method for the unsupervised classification of cells within a sample into a number of clusters and in the case of multiple samples, the matching of these clusters across the samples. The two tasks of clustering and matching of the clusters are performed simultaneously within the JCM framework. In this paper, we consider the case where there is a number of distinct classes of samples whose class of origin is known, and the problem is to classify a new sample of unknown class of origin to one of these predefined classes. For example, the different classes might correspond to the types of a particular disease or to the various health outcomes of a patient subsequent to a course of treatment. We show and demonstrate on some real datasets how the JCM procedure can be used to carry out this supervised classification task. A mixture distribution is used to model the distribution of the expressions of a fixed set of markers for each cell in a sample with the components in the mixture model corresponding to the various populations of cells in the composition of the sample. For each class of samples, a class template is formed by the adoption of random-effects terms to model the inter-sample variation within a class. The classification of a new unclassified sample is undertaken by assigning the unclassified sample to the class that minimizes the Kullback-Leibler distance between its fitted mixture density and each class density provided by the class templates.

研究の動機と目的

  • 新しいフローサイトメトリー・サンプルを事前に定義された疾患または健康状態のクラスに分類する課題に対処すること。
  • 高次元で複雑なフローサイトメトリー・データにおいて、手動ゲーティングや特徴ベースの分類器の限界を克服すること。
  • 要約統計量ではなく、密度の完全な情報を活用するパラメトリックでモデルベースのアプローチを用いて、サンプル分類を実現すること。
  • 実世界のデータセット上で、JCMに基づく分類器の性能を、既存の最先端手法と比較して評価すること。

提案手法

  • 細胞集団を、制限付き歪み正規分布やt分布などの有限混合モデルで表現し、非対称性と厚い尾部を捉える。
  • JCMフレームワークを用いて、同時にクラスタリングとサンプル間クラスターマッチングを実行し、複数のサンプル間で細胞集団を整合させる。
  • 階層的混合モデルにおけるランダム効果項を用いてサンプル間の変動をモデル化することで、各事前に定義されたクラスのテンプレートを構築する。
  • 新しい未分類サンプルを、その適合された混合密度と各クラステンプレートの密度との間のKullback-Leibler(KL)発散を計算することで分類する。
  • EMアルゴリズムを用いて、成分の平均、共分散、混合割合、ランダム効果項などのモデルパラメータを推定する。
  • 完全なモデル適合の前段階として、次元削減(PCA、NMF、またはGMF)または外れ値除去(マーカー部分集合に対するJCMを用いて)を実施し、ロバストネスとスケーラビリティを向上させる。

実験結果

リサーチクエスチョン

  • RQ1JCM手順は、事前に定義されたクラスに分類するフローサイトメトリー・サンプルに効果的に適応可能か?
  • RQ2JCMに基づく分類の性能は、SVM、Citrus、HDPGMMなどの既存の特徴ベース分類器と比べてどうか?
  • RQ3KL発散を用いた密度ベースの比較が、限られた統計量(例:クラスタープロポーション)に依存する手法よりも分類精度を向上させるか?
  • RQ4臨床的フローサイトメトリー・データセットにおける高次元データとサンプル間変動に対して、JCMアプローチはどれほどロバストか?

主な発見

  • AML分類チャレンジにおいて、JCMは受受曲線下積分(AUC)がほぼ1.0に近く、比較された6つの手法の中で最高であった。
  • JCMはすべてのAMLサンプルを正しく分類し、感度が1.0を達成したが、他の手法では0.95を上回る感度を示したものはなかった。
  • BCRデータセットでは、FスコアとAUCの両方でJCMが他のすべての手法を上回り、優れた分類精度を示した。
  • 完全なパラメトリック密度モデルとKL発散を用いた比較により、限られた統計量に依存する特徴ベース手法よりも、より正確な分類が可能になった。
  • クラステンプレートにランダム効果項を統合することで、同じクラス内でのサンプル間の変動を効果的に捉え、一般化性能が向上した。
  • 外れ値除去や次元削減などの前処理を経て、高次元設定下でも、この手法はロバストな性能を維持した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。