Skip to main content
QUICK REVIEW

[論文レビュー] Edge principal components and squash clustering: using the special structure of phylogenetic placement data for sample comparison

Frederick A. Matsen, Steven N. Evans|arXiv (Cornell University)|Jul 25, 2011
Genomics and Phylogenetic Studies被引用数 8
ひとこと要約

本稿では、配置データの系統的構造を活用して、マイクロバイオームサンプル比較における解釈可能性を向上させる2つの手法、Edge PCA と Squash Clustering を提案する。Edge PCA は主成分を参照系統樹上の重み付きエッジとして可視化するのに対し、Squash Clustering はエッジ長が平均化された微生物集団間の意味のある距離を反映する根付き木を生成し、UPGMA よりも明確な生物学的解釈が可能になる。

ABSTRACT

Principal components (PCA) and hierarchical clustering are two of the most heavily used techniques for analyzing the differences between nucleic acid sequence samples sampled from a given environment. However, a classical application of these techniques to distances computed between samples can lack transparency because there is no ready interpretation of the axes of classical PCA plots, and it is difficult to assign any clear intuitive meaning to either the internal nodes or the edge lengths of trees produced by distance-based hierarchical clustering methods such as UPGMA. We show that more interesting and interpretable results are produced by two new methods that leverage the special structure of phylogenetic placement data. Edge principal components analysis enables the detection of important differences between samples that contain closely related taxa. Each principal component axis is simply a collection of signed weights on the edges of the phylogenetic tree, and these weights are easily visualized by a suitable thickening and coloring of the edges. Squash clustering outputs a (rooted) clustering tree in which each internal node corresponds to an appropriate "average" of the original samples at the leaves below the node. Moreover, the length of an edge is a suitably defined distance between the averaged samples associated with the two incident nodes, rather than the less interpretable average of distances produced by UPGMA. We present these methods and illustrate their use with data from the microbiome of the human vagina.

研究の動機と目的

  • マイクロバイオームデータにおけるUniFrac距離に古典的手法のPCA や階層的クラスタリングを適用した際の解釈可能性の欠如に対処すること。
  • 系統的配置データの系統的構造を明示的に活用して、サンプル比較に適した手法を開発すること。
  • 主成分を参照系統樹上の特定エッジに基づくものとして可視化・解釈可能にするための方法を可能にすること。
  • エッジ長が平均化された微生物集団間の生物学的に意味のある距離に対応する階層的クラスタリング手法を構築すること。
  • 特に近縁種に対して、微生物集団の差異を分析する際の透明性と生物学的洞察を向上させること。

提案手法

  • Edge PCA は、参照系統樹の内部エッジにおける配置割合の差異に基づいて主成分を計算する。
  • 各主成分軸は、木のエッジに符号付き重みの集合として表現され、エッジの太さと色で可視化される。
  • Squash Clustering は、参照樹上の系統的配置分布を組み込んだ新しいクラスタ間距離定義を用いる。
  • この手法は、各内部ノードが平均化された微生物集団分布を表し、エッジ長がそれらの分布間の距離を反映する根付きクラスタリング木を構築する。
  • アルゴリズムは、再構築可能性パラメータが子クラスタ間の類似性を制御するように、参照木の効果的分割に基づいて再帰的にクラスタを分割する。
  • シミュレーションでは、合成配置データの生成のため、カット数をポアソン分布、サブセット割り当てを二項分布でモデル化する。

実験結果

リサーチクエスチョン

  • RQ1配置データの系統的構造を活用することで、マイクロバイオームデータの主成分分析を、古典的手法よりも解釈可能にできるか?
  • RQ2エッジに基づく主成分は、古典的手法のPCA よりも、近縁種を含むわずかだが一貫したサンプル差をより効果的に検出できるか?
  • RQ3階層的クラスタリングを再定義し、エッジ長が平均化された微生物集団間の生物学的に意味のある距離に対応する木を生成できるか?
  • RQ4シミュレーテッドデータからクラスタリング関係を再構築する際、Squash Clustering はUPGMA よりも真の木のトポロジーをどれほどよく保持するか?
  • RQ5再構築可能性パラメータは、クラスタリング木の再構築精度と類似性にどの程度影響を及ぼすか?

主な発見

  • Edge PCA は、近縁種を含むわずかだが一貫した差を効果的に検出でき、古典的手法のPCA が見逃す可能性がある。
  • Edge PCA の主成分軸は、参照系統樹上の特定エッジの重み付き寄与として直接解釈可能であり、視覚的・生物学的解釈が可能である。
  • Squash Clustering は、エッジ長が平均化された微生物集団分布間の意味のある距離に対応する根付きクラスタリング木を生成するが、UPGMA とは異なり、より解釈可能な平均距離を用いる。
  • この手法は、クラスタリング木の各ノードに、参照木上での自然な質量分布を割り当て、生物学的解釈性を向上させる。
  • シミュレーションでは、再構築可能性パラメータ $r_t$ を高く設定したSquash Clustering は、根付きRobinson-Foulds距離で測定したところ、真の木に近いクラスタリング木を生成した。
  • 6葉の木における根付きRobinson-Foulds距離の最大値は4であり、この手法はパラメータ設定の変化に敏感に反応し、木のトポロジー再構築に影響を及げた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。