[Paper Review] Sample Size Dependent Species Models
This paper introduces a sample size dependent species model based on a generalized negative binomial process that enables flexible, nonparametric Bayesian estimation of Simpson’s index of diversity. By modeling cluster structures dependent on sample size, the approach improves accuracy in estimating species evenness from sequencing count data, outperforming traditional exchangeable partition models in simulation and real-world genomic and immunological datasets.
Motivated by the fundamental problem of measuring species diversity, this paper introduces the concept of a cluster structure to define an exchangeable cluster probability function that governs the joint distribution of a random count and its exchangeable random partitions. A cluster structure, naturally arising from a completely random measure mixed Poisson process, allows the probability distribution of the random partitions of a subset of a sample to be dependent on the sample size, a distinct and motivated feature that differs it from a partition structure. A generalized negative binomial process model is proposed to generate a cluster structure, where in the prior the number of clusters is finite and Poisson distributed, and the cluster sizes follow a truncated negative binomial distribution. We construct a nonparametric Bayesian estimator of Simpson's index of diversity under the generalized negative binomial process. We illustrate our results through the analysis of two real sequencing count datasets.
Motivation & Objective
- To address the limitation of existing exchangeable partition models that assume sample size independence in cluster probability functions.
- To develop a flexible, nonparametric Bayesian framework for modeling species diversity that accounts for sample size effects in species abundance data.
- To construct a robust estimator of Simpson’s index of diversity that improves upon traditional frequentist estimates in small-sample and high-diversity settings.
- To enable accurate inference of species evenness from count data in genomics and immunology, such as T-cell receptor and EST sequencing datasets.
- To provide a foundation for modeling random count vectors and latent count matrices in Bayesian nonparametric mixture models.
Proposed method
- Proposes a cluster structure derived from a completely random measure mixed Poisson process, allowing partition probabilities to depend on sample size.
- Introduces a generalized negative binomial process (GNBP) with a finite, Poisson-distributed number of clusters and truncated negative binomial cluster sizes.
- Derives an exchangeable cluster probability function (ECPF) that generalizes the EPPF by incorporating sample size dependence.
- Develops a nonparametric Bayesian estimator of Simpson’s index using the GNBP prior, enabling posterior inference via MCMC.
- Applies the Chinese restaurant process-like sampling rule to generate partitions with size-dependent predictive probabilities.
- Uses MCMC with Gibbs sampling to estimate posterior distributions over parameters (γ₀, a, p) and compute posterior predictive diversity estimates.
Experimental results
Research questions
- RQ1Can a species model be constructed such that the distribution of random partitions depends explicitly on the sample size, rather than being invariant as in partition structures?
- RQ2How can a nonparametric Bayesian model be designed to estimate Simpson’s index of diversity with improved accuracy in small-sample and high-diversity settings?
- RQ3What is the performance of a generalized negative binomial process in modeling species abundance frequency counts compared to standard EPPF-based models?
- RQ4Can the proposed model effectively capture species evenness in real-world sequencing datasets, such as T-cell receptor and EST data?
- RQ5How does allowing the discount parameter a to be inferred rather than fixed affect estimation accuracy of Simpson’s index?
Key findings
- The proposed generalized negative binomial process model with inferred discount parameter a < 1 achieved 99% 95% coverage and 69% 50% coverage for the true Simpson’s index in EST data simulations.
- When a was restricted to 0 ≤ a < 1, the method achieved a mean bias of 0.48 × 10⁻³ and median bias of 1.11 × 10⁻³, significantly outperforming fixed a = 0.5 or a = 0.
- The model demonstrated superior performance on both T-cell receptor and EST datasets, showing lower Simpson’s index in regulatory T-cells of diabetic mice, indicating reduced diversity.
- The simulation study showed that fixing a to -1 or 0 resulted in 0% coverage for both 50% and 95% credible intervals, highlighting the importance of allowing a to be inferred.
- The model’s ability to incorporate sample size dependence led to more accurate posterior estimates of diversity, especially in high-diversity, low-sample-size scenarios.
- The framework enables efficient MCMC inference and provides a foundation for modeling random count matrices and latent cluster structures in Bayesian nonparametric models.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.