[Paper Review] Beta-Negative Binomial Process and Poisson Factor Analysis
This paper proposes a beta-negative binomial (BNB) process as a nonparametric Bayesian prior for infinite Poisson factor analysis (PFA), enabling flexible modeling of overdispersed count data through a beta-gamma-gamma-Poisson hierarchical structure. The method automatically infers the number of active factors and achieves state-of-the-art performance in document count matrix factorization, with lower perplexity than existing models.
A beta-negative binomial (BNB) process is proposed, leading to a beta-gamma-Poisson process, which may be viewed as a "multi-scoop" generalization of the beta-Bernoulli process. The BNB process is augmented into a beta-gamma-gamma-Poisson hierarchical structure, and applied as a nonparametric Bayesian prior for an infinite Poisson factor analysis model. A finite approximation for the beta process Levy random measure is constructed for convenient implementation. Efficient MCMC computations are performed with data augmentation and marginalization techniques. Encouraging results are shown on document count matrix factorization.
Motivation & Objective
- To address the limitations of Gaussian-based latent factor models in modeling discrete, nonnegative, overdispersed count data.
- To develop a nonparametric Bayesian prior that allows flexible modeling of both mean and variance in latent count structures.
- To extend the beta-Bernoulli process to a 'multi-scoop' generalization using negative binomially distributed counts.
- To enable efficient inference for infinite-dimensional count matrix factorization with automatic model selection.
- To improve topic modeling performance by capturing diverse topic characteristics through learned negative binomial parameters.
Proposed method
- Proposes a beta-negative binomial (BNB) process by extending the beta process to a marked space $[0,1] \times \mathbb{R}^+ \times \Omega$, allowing for count-valued marks.
- Constructs a hierarchical beta-gamma-gamma-Poisson process, where the Poisson rate is drawn from a gamma distribution, and the gamma rate is drawn from a beta process.
- Uses a finite approximation of the beta process Lévy random measure to enable practical MCMC implementation.
- Employs data augmentation and marginalization techniques to perform efficient MCMC inference over the joint posterior distribution.
- Leverages conjugate relationships between beta, gamma, Poisson, negative binomial, and Dirichlet distributions to simplify computation.
- Applies the model to document count matrix factorization, interpreting latent factors as topics with count-based contributions.
Experimental results
Research questions
- RQ1Can a nonparametric Bayesian prior be developed to model multivariate count data with flexible overdispersion?
- RQ2How can the beta-Bernoulli process be generalized to allow for multiple counts per latent feature, rather than binary presence/absence?
- RQ3Can a hierarchical structure be designed to jointly learn the mean and variance of latent count factors in a nonparametric way?
- RQ4Does the proposed model outperform existing PFA and topic models in terms of held-out perplexity on document corpora?
- RQ5How do the inferred negative binomial parameters ($r_k$, $p_k$) relate to topic interpretability and characteristics?
Key findings
- The proposed $\beta\gamma\Gamma$-PFA model achieves the lowest held-out perplexity on both JACM and PsyRev document corpora, outperforming $\Gamma$-PFA, Dirich-PFA, $\beta\Gamma$-PFA, and $S\gamma\Gamma$-PFA.
- For the JACM corpus, $\beta\gamma\Gamma$-PFA inferred 132 active factors at the final MCMC iteration, while for PsyRev it inferred 209, demonstrating automatic model selection.
- Smaller values of the topic Dirichlet prior $a_\phi = 0.01$ led to higher inferred factor counts and better predictive performance, though too small values caused over-specialization.
- The model effectively absorbs stop words into a few dominant topics with large mean and small variance, preserving interpretability of other topics.
- Topics with large mean and large variance (e.g., 'rivalry, binocular, monocular') or small mean and large variance (e.g., 'search, binary, tree') were accurately captured through learned $r_k$ and $p_k$ parameters.
- The hierarchical structure enables joint learning of factor loadings and their overdispersion, leading to more robust and interpretable topic models.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.