[論文レビュー] Bayesian Modeling of Microbiome Data for Differential Abundance Analysis
本稿では、マイクロバイオームデータにおける差分豊度解析のためのベイジアン階層モデルZINB-DPPを提案する。この手法は、ゼロインflatedネガティブバイノミアル(ZINB)モデルとディリクレ過程事前分布を統合し、特徴選択および系統発生的構造を組み込む。本手法は、過剰なゼロ、分散拡大、不均一なシークエンシング深度を示す生物学的に関連のある分離株を検出する能力において、既存手法を上回り、誤発見率(FDR)を制御するとともに進化的関係を組み込む。
The advances of next-generation sequencing technology have accelerated study of the microbiome and stimulated the high throughput profiling of metagenomes. The large volume of sequenced data has encouraged the rise of various studies for detecting differentially abundant taxonomic features across healthy and diseased populations, with the ultimate goal of deciphering the relationship between the microbiome diversity and health conditions. As the microbiome data are high-dimensional, typically featuring by uneven sampling depth, overdispersion and a huge amount of zeros, these data characteristics often hamper the downstream analysis. Moreover, the taxonomic features are implicitly imposed by the phylogenetic tree structure and often ignored. To overcome these challenges, we propose a Bayesian hierarchical modeling framework for the analysis of microbiome count data for differential abundance analysis. Under this framework, we introduce a bi-level Bayesian hierarchical model that allows a flexible choice of the count generating process, and hyperpriors in the feature selection scheme. We particularly focus on employing a zero-inflated negative binomial model with a Bayesian nonparametric prior model on the bottom level, and applying Gaussian mixture models for differentially abundant taxa detection on the top level. Our method allows for the simultaneous modeling of sample heterogeneity and detecting differentially abundant taxa. We conducted comprehensive simulations and summarized the improved statistical performances of the proposed model. We applied the model in two real microbiome study datasets and successfully identified biologically validated differentially abundant taxa. We hope that the proposed framework and model can facilitate further microbiome studies and elucidate disease etiology.
研究の動機と目的
- 高次元マイクロバイオームカウントデータにおけるゼロ過剰、分散拡大、不均一なシークエンシング深度の課題に対処すること。
- 情報的事前分布によるモデルベース正規化を統合し、恣意的な事前正規化を回避する統合的統計フレームワークの構築。
- 疾患状態(例:大腸癌、統合失調症)間での差分豊度を有する分離株を同定する際の検出力と正確性の向上。
- マルコフ確率場事前分布を用いて系統発生的関係を差分豊度検定に組み込むこと。
- ベイジアンFDR推定を用いた誤発見率制御により、結果の再現性と生物学的妥当性の向上。
提案手法
- モデルは二段階の階層構造を採用:下位レベルでゼロ過剰および分散拡大を示すマイクロバイオームカウントデータをZINB分布でモデル化。
- モデルベース正規化は、事前分布における確率的制約によって達成され、事前正規化の必要性がなくなる。
- 上位レベルの混合正規分布にディリクレ過程事前分布を適用することで、非パラメトリックな特徴選択と差分豊度を示す分離株の自動特定が可能になる。
- 系統発生的構造は、木の隣接する分離株が差分豊度状態を共有するのを促進するマルコフ確率場事前分布を用いて組み込まれる。
- すべてのパラメータはマルコフ連鎖モンテカルロ(MCMC)サンプリングにより推定され、完全な後方分布推論が可能になる。
- ベイジアン誤発見率(FDR)制御が適用され、後方包含確率の不確実性を考慮した有意な分離株の同定が可能になる。
実験結果
リサーチクエスチョン
- RQ1情報的事前分布による正規化を必要とせず、ゼロ過剰、分散拡大、高次元マイクロバイオームカウントデータを効果的に処理できるベイジアン階層モデルは存在するか?
- RQ2マルコフ確率場事前分布による系統発生的構造の組み込みは、生物学的に意味のある差分豊度を示す分離株の検出をどのように向上させるか?
- RQ3従来の手法(例:Kruskal–Wallis、DESeq2、edgeR、metagenomeSeq)と比較して、ZINB-DPPモデルは統計的検出力およびFDR制御においてどのように異なるか?
- RQ4大腸癌や統合失調症などの実世界のデータセットにおいて、本モデルは既知のマイクロバイオーム-疾患関連性をどの程度回復できるか?
- RQ5Fusobacterium nucleatumとCampylobacterのような生物学的に関連するが、標準的手法で見逃されがちな共発生分離株を同定できるか?
主な発見
- 大腸癌データセットでは、ZINB-DPPモデルが1%のベイジアンFDR閾値で10種の差分豊度を示す分離株を同定し、そのうち7つが既存の生物学的証拠で裏付けられていた。
- ZINB-DPPモデルは、CRCにおいてSynergistaceaeからSynergistetesに至る系統発生的系統が豊富であることを的確に同定し、先行研究で確認された結果を再現した。
- Kruskal–Wallisと比較して、12種の分離株を報告したが、そのうち生物学的裏付けのあるのは7種にとどまった。一方、ZINB-DPPモデルは11種のうち6種が文献で確認されており、より高い正確性を示した。
- 統合失調症研究において、ZINB-DPPは5%のベイジアンFDRで8種の差分豊度を示す分離株を同定し、そのうち5つがDESeq2およびmetagenomeSeqの結果と重複していた。また、Veillonella parvulaの独自同定も達成した。
- FDR制御の観点から、ZINB-DPPモデルはDESeq2やedgeRを上回り、後者よりも誤検出数が少ないことが実証された。
- ZINB-DPPモデルは、Fusobacterium nucleatumとCampylobacterの共発生パターンを同定できたが、metagenomeSeqやDMモデルでは同定されなかった。これは、生物学的共発生に対する感受性の向上を示している。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。