[Paper Review] Modeling population structure under hierarchical Dirichlet processes
This paper proposes a Bayesian nonparametric model based on the Hierarchical Dirichlet Process (HDP) to infer population admixture while accounting for linkage disequilibrium (LD) between loci. By modeling correlated ancestry across adjacent genetic markers, the method improves upon traditional approaches like STRUCTURE by allowing for variable-length ancestral segments and nonparametric inference on the number of ancestral populations, demonstrating accurate detection of rare haplotypes in real data.
We propose a Bayesian nonparametric model to infer population admixture, extending the Hierarchical Dirichlet Process to allow for correlation between loci due to Linkage Disequilibrium. Given multilocus genotype data from a sample of individuals, the model allows inferring classifying individuals as unadmixed or admixed, inferring the number of subpopulations ancestral to an admixed population and the population of origin of chromosomal regions. Our model does not assume any specific mutation process and can be applied to most of the commonly used genetic markers. We present a MCMC algorithm to perform posterior inference from the model and discuss methods to summarise the MCMC output for the analysis of population admixture. We demonstrate the performance of the proposed model in simulations and in a real application, using genetic data from the EDAR gene, which is considered to be ancestry-informative due to well-known variations in allele frequency as well as phenotypic effects across ancestry. The structure analysis of this dataset leads to the identification of a rare haplotype in Europeans.
Motivation & Objective
- To address the limitation of existing admixture models in handling linkage disequilibrium (LD) between closely spaced genetic loci.
- To develop a method that infers the number of ancestral populations nonparametrically, without assuming a fixed K.
- To enable accurate assignment of chromosomal segments to ancestral populations, reflecting the mosaic nature of admixed genomes.
- To model population structure without assuming a specific mutation process, making it applicable to diverse genetic markers.
- To provide a fully Bayesian framework with MCMC inference and posterior summarization tools for population admixture analysis.
Proposed method
- Extends the Hierarchical Dirichlet Process (HDP) to model correlated ancestry states across loci, capturing LD through a Chinese Restaurant Process (CRP) representation of population assignments.
- Uses a Gibbs sampler with conditional updates for population memberships (z_il), ancestry proportions (n_ik, m_ik), allele frequencies (θ_kl), and hyperparameters (α, α₀, r, μ_l).
- Incorporates a Beta-Binomial conjugate structure for SNP data, where θ_kl follows a Beta distribution conditional on observed genotypes in each population.
- Employs auxiliary variables and Polya-Gamma data augmentation for efficient posterior updates of concentration parameters α and α₀.
- Applies a random walk Metropolis step for updating the Poisson process rate r and hyperparameters μ_l in the base measure H.
- Uses a nonparametric prior on the base measure H to allow flexible modeling of ancestral allele frequencies without assuming a fixed parametric form.
Experimental results
Research questions
- RQ1How can population admixture be modeled when loci are correlated due to linkage disequilibrium?
- RQ2Can a Bayesian nonparametric model infer the number of ancestral populations without pre-specifying K?
- RQ3To what extent does accounting for LD improve the accuracy of chromosomal ancestry inference?
- RQ4How can MCMC sampling be designed to efficiently explore complex, high-dimensional posterior distributions in admixture models?
- RQ5What insights can be gained about rare genetic variants, such as the EDAR haplotype, using this model?
Key findings
- The model successfully detects a rare haplotype in Europeans within the EDAR gene, which was previously overlooked by standard methods.
- Posterior inference on the number of ancestral populations is nonparametrically determined, with the model automatically adapting to the data-driven number of clusters.
- The inclusion of LD modeling leads to more accurate inference of ancestral segment lengths and population of origin for chromosomal regions.
- Simulations show that the model outperforms standard HDP and STRUCTURE in recovering true population structure under LD conditions.
- The MCMC algorithm converges well and provides reliable posterior summaries for ancestry proportions and population assignments.
- The method is robust across different genetic markers and does not require assumptions about the underlying mutation process.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.