[논문 리뷰] Bayesian Modeling of Microbiome Data for Differential Abundance Analysis
이 논문은 제로-과잉, 과분산, 비균형 시퀀싱 깊이를 포함한 고차원 미생물군집 카운트 데이터에서 차별적 번성 분석을 위한 베이지안 계층 모델인 ZINB-DPP를 제안한다. 이 모델은 제로-과잉 음수 이항(ZINB) 모델링과 딜리클 프로세스 사전분포를 융합하여 특징 선택과 계통발생적 구조를 통합한다. 이 방법은 생물학적으로 관련성이 있는 분류군, 특히 제로가 과도하게 포함된 경우, 과분산, 비균형 시퀀싱 깊이를 가진 경우에도 기존 방법보다 뛰어난 성능을 보이며, 거짓 발견률(FDR)을 제어하고 진화적 관계를 통합한다.
The advances of next-generation sequencing technology have accelerated study of the microbiome and stimulated the high throughput profiling of metagenomes. The large volume of sequenced data has encouraged the rise of various studies for detecting differentially abundant taxonomic features across healthy and diseased populations, with the ultimate goal of deciphering the relationship between the microbiome diversity and health conditions. As the microbiome data are high-dimensional, typically featuring by uneven sampling depth, overdispersion and a huge amount of zeros, these data characteristics often hamper the downstream analysis. Moreover, the taxonomic features are implicitly imposed by the phylogenetic tree structure and often ignored. To overcome these challenges, we propose a Bayesian hierarchical modeling framework for the analysis of microbiome count data for differential abundance analysis. Under this framework, we introduce a bi-level Bayesian hierarchical model that allows a flexible choice of the count generating process, and hyperpriors in the feature selection scheme. We particularly focus on employing a zero-inflated negative binomial model with a Bayesian nonparametric prior model on the bottom level, and applying Gaussian mixture models for differentially abundant taxa detection on the top level. Our method allows for the simultaneous modeling of sample heterogeneity and detecting differentially abundant taxa. We conducted comprehensive simulations and summarized the improved statistical performances of the proposed model. We applied the model in two real microbiome study datasets and successfully identified biologically validated differentially abundant taxa. We hope that the proposed framework and model can facilitate further microbiome studies and elucidate disease etiology.
연구 동기 및 목표
- 고차원의 미생물군집 카운트 데이터에서 제로 과잉, 과분산, 비균형 시퀀싱 깊이 문제를 해결하기 위해.
- 정보성 사전분포를 통한 모델 기반 정규화를 통합하여 즉흥적인 사전 정규화가 필요 없는 통합 통계 프레임워크를 개발하기 위해.
- 질병 상태(예: 대장직장암, 조현병) 간에 차별적 번성이 높은 분류군을 탐지하는 데 있어 검출 능력과 정확도를 향상시키기 위해.
- 마르코프 무작위 필드 사전분포를 사용하여 계통발생적 관계를 차별적 번성 검정에 통합하기 위해.
- 베이지안 FDR 추정을 통해 거짓 발견률을 제어하여 결과의 재현성과 생물학적 관련성을 향상시키기 위해.
제안 방법
- 모델은 이중 계층 계층적 구조를 사용한다: 하위 수준에서 제로 과잉과 과분산을 가진 미생물군집 카운트 데이터를 모델링하기 위해 ZINB 분포를 적용한다.
- 모델 기반 정규화는 사전분포의 스토케스틱 제약 조건을 통해 달성되며, 사전 정규화가 필요 없어진다.
- 상위 수준에서 딜리클 프로세스 사전분포를 갖는 정규분포 혼합 모델은 비모수적 특징 선택과 자동으로 차별적 번성이 높은 분류군을 결정하는 데 기여한다.
- 계통발생적 구조는 마르코프 무작위 필드 사전분포를 사용하여 트리상의 인접한 분류군이 차별적 번성 상태를 공유하도록 유도한다.
- 모든 매개변수는 마르코프 체인 몬테카를로(MCMC) 샘플링을 통해 추정되며, 전체 사후 추론이 가능하다.
- 베이지안 거짓 발견률(FDR) 제어는 사후 포함 확률의 불확실성을 고려하여 유의미한 분류군을 식별하는 데 적용된다.
실험 결과
연구 질문
- RQ1정보성 사전분포를 통한 사전 정규화 없이도, 제로 과잉, 과분산, 고차원의 미생물군집 카운트 데이터를 효과적으로 다룰 수 있는 베이지안 계층 모델이 존재하는가?
- RQ2마르코프 무작위 필드 사전분포를 통한 계통발생적 구조의 통합이 생물학적으로 의미 있는 차별적 번성이 높은 분류군 탐지에 어떻게 기여하는가?
- RQ3기존 방법(예: 크러스칼-월리스, DESeq2, edgeR, metagenomeSeq)과 비교했을 때 ZINB-DPP 모델의 통계적 검출 능력과 FDR 제어 능력은 어떠한가?
- RQ4실제 데이터셋(대장직장암, 조현병 등)에서 모델이 기존에 알려진 미생물군집-질병 연관성을 얼마나 잘 복원하는가?
- RQ5표준 방법에서 놓칠 수 있는 생물학적으로 연관된 공생 분류군(예: Fusobacterium nucleatum 및 Campylobacter)을 탐지할 수 있는가?
주요 결과
- 대장직장암 데이터셋에서 ZINB-DPP 모델은 1% 베이지안 FDR 기준으로 10개의 차별적 번성이 높은 종을 탐지했으며, 그 중 7개는 이전 생물학적 증거에 의해 지지되었다.
- ZINB-DPP 모델은 CRC에서 Synergistaceae에서 Synergistetes 계통으로의 부풀림을 성공적으로 식별했으며, 이는 이전 연구에서 검증된 결과였다.
- 크러스칼-월리스 방법은 12개의 종을 보고했지만, 그 중 7개만 생물학적으로 지지되었고, ZINB-DPP 모델은 문헌에서 확인된 6/11개 종을 통해 더 높은 정밀도를 달성했다.
- 조현병 연구에서 ZINB-DPP는 5% 베이지안 FDR 기준으로 8개의 차별적 번성이 높은 분류군을 식별했으며, 그 중 5개는 DESeq2와 metagenomeSeq의 결과와 겹쳤고, Veillonella parvula의 유일한 탐지 결과도 있었다.
- DESeq2와 edgeR에 비해 ZINB-DPP 모델이 FDR 제어에서 뛰어나, 후자보다 더 많은 탐지 수에도 불구하고 더 적은 거짓 양성 결과를 보였다.
- ZINB-DPP 모델은 Fusobacterium nucleatum과 Campylobacter 간의 공생 패턴을 탐지했으며, 이는 metagenomeSeq와 DM 모델에서 놓친 바이오로지컬 공생성을 반영하여 민감도 향상을 보였다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.