[論文レビュー] Clustering transformed compositional data using K-means, with applications in gene expression and bicycle sharing system data
本稿では、遺伝子発現データおよび自転車共有データにおいて、クラスタリング性能を向上させるために、データ変換を用いたK-meansクラスタリング戦略を提案する。具体的には、センター・ログ比率(CLR)および新規のログセンター・ログ比率(logCLR)を用いる。変換を施したデータは、ゼロ値または近似ゼロ値をより適切に扱い、非漸近的ペナルティ付き基準とスロープヒューリスティクスを用いて最適なクラスタ数を特定することで、元の未変換データに比べて優れた性能を発揮する。
Although there is no shortage of clustering algorithms proposed in the literature, the question of the most relevant strategy for clustering compositional data (i.e., data made up of profiles, whose rows belong to the simplex) remains largely unexplored in cases where the observed value of an observation is equal or close to zero for one or more samples. This work is motivated by the analysis of two sets of compositional data, both focused on the categorization of profiles but arising from considerably different applications: (1) identifying groups of co-expressed genes from high-throughput RNA sequencing data, in which a given gene may be completely silent in one or more experimental conditions; and (2) finding patterns in the usage of stations over the course of one week in the Velib' bicycle sharing system in Paris, France. For both of these applications, we focus on the use of appropriately chosen data transformations, including the Centered Log Ratio and a novel extension we propose called the Log Centered Log Ratio, in conjunction with the K-means algorithm. We use a nonasymptotic penalized criterion, whose penalty is calibrated with the slope heuristics, to select the number of clusters present in the data. Finally, we illustrate the performance of this clustering strategy, which is implemented in the Bioconductor package coseq, on both the gene expression and bicycle sharing system data.
研究の動機と目的
- ゼロまたは近似ゼロの値を含む組成データに対して、体系的なクラスタリング戦略が不足している問題に対処すること。
- データ変換が組成データ設定下でのK-meansクラスタリング性能に与える影響を調査すること。
- 単体の頂点付近に位置する極端な観測値の分離を向上させるために、新規の変換であるログセンター・ログ比率(logCLR)を提案すること。
- 非漸近的ペナルティ付き基準をスロープヒューリスティクスで補正し、最適なクラスタ数を特定すること。
- 実世界のデータセット(RNA-seq遺伝子発現データおよびVelib’自転車共有システムデータ)を用いて、本手法の有効性を実証すること。
提案手法
- 未変換の組成データ、CLR変換済みデータ、logCLR変換済みデータの3種類のデータ形式に対して、ユークリッド距離を用いたK-meansクラスタリングを適用する。
- 近似ゼロ成分を有するプロファイルの分離を改善するため、CLRの変更版としてlogCLR変換を導入する。
- 過学習を回避するため、非漸近的ペナルティ付き基準(例:BICに類似)にスロープヒューリスティクスを適用し、クラスタ数を決定する。
- Bioconductorパッケージcoseqを用いて、組成データの再現可能解析のための完全なパイプラインを実装する。
- プロファイル可視化および空間マッピング(例:パリの地理的駅クラスタ)を用いてクラスタリング結果を評価する。
- 変換手法ごとの結果を比較し、生物学的または行動的に意味のあるグループ化を最もよく捉えている手法を特定する。

実験結果
リサーチクエスチョン
- RQ1未変換、CLR、logCLRの異なるデータ変換が、ゼロまたは近似ゼロの値を含む組成データにおけるK-meansクラスタリング性能にどのように影響を与えるか?
- RQ2提案されたlogCLR変換は、標準のCLRに比べて、極端に特異的またはエッジ・ケースの組成プロファイルをよりよく捉えることができるか?
- RQ3真の構造にスパースまたは単体の頂点付近に位置するプロファイルを含む組成データにおいて、最適なクラスタ数は何か?
- RQ4スロープヒューリスティクスに基づくペナルティ付き基準は、実世界の組成データセットにおいてクラスタ数を適切に選択できるか?
- RQ5クラスタリング結果は、遺伝子共発現や都市移動行動といった現実世界のパターンとして、どの程度解釈可能か?
主な発見
- logCLR変換は、RNA-seqおよびVelib’自転車共有データの両方において、最も満足のいくクラスタリング結果をもたらした。特に、行動的に一貫性のあるユーザーまたは遺伝子プロファイルを明確に特定できた。
- Velib’データからのクラスタプロファイルは、既知の都市地域と強く空間的整合性を示した。例えば、都市中心部に位置するクラスタ3、住宅地に位置するクラスタ7、および高標高V+ステーションで夜間の再配分が必要なクラスタ11。
- logCLR変換を用いたK-meansアルゴリズムは、一部の遺伝子が特定の条件下で無発現であっても、生物学的に意味のある共発現グループを適切に同定できた。
- 非漸近的ペナルティ付き基準にスロープヒューリスティクスを適用した手法は、過学習を回避し、安定したモデル選択を可能にした。
- 本手法の結果は、原始カウントにカイ二乗距離を用いた場合と同等の性能を示し、CLR/logCLR変換が組成データクラスタリングの有効な代替手法であることを示唆した。
- logCLR変換は、近似ゼロ成分を有するプロファイルの分離において特に効果的であり、エッジ・ケース検出において、未変換および標準CLRアプローチを上回った。

より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。