[Paper Review] Stochastic blockmodels for exchangeable collections of networks
This paper introduces a Bayesian nonparametric stochastic blockmodel for jointly analyzing exchangeable collections of networks, enabling simultaneous estimation of community structures across multiple networks while allowing for uncertainty quantification and model comparison. The method employs a Chinese restaurant process prior over partitions and uses MCMC with split-merge moves to improve mixing and convergence, yielding interpretable, data-driven community detection across diverse network types.
We construct a novel class of stochastic blockmodels using Bayesian nonparametric mixtures. These model allows us to jointly estimate the structure of multiple networks and explicitly compare the community structures underlying them, while allowing us to capture realistic properties of the underlying networks. Inference is carried out using MCMC algorithms that incorporates sequentially allocated split-merge steps to improve mixing. The models are illustrated using a simulation study and a variety of real-life examples.
Motivation & Objective
- To develop a statistical model that jointly estimates community structures across multiple exchangeable networks, capturing shared and distinct structural patterns.
- To enable direct comparison of community structures across networks while accounting for uncertainty in network topology and group assignments.
- To extend stochastic blockmodels to infinite-dimensional settings where the number of communities is inferred from data, avoiding model selection bias.
- To incorporate diverse network types—directed, undirected, binary, and count-valued—within a unified hierarchical framework.
- To improve posterior sampling efficiency through sequential split-merge MCMC moves that enhance mixing in high-dimensional partition spaces.
Proposed method
- Uses a Chinese restaurant process prior over partitions to allow for an unknown and potentially infinite number of communities in each network.
- Employs a hierarchical Bayesian model where network-specific parameters are shared across actors via group-level distributions.
- Applies MCMC with sequentially allocated split-merge moves to improve mixing in the space of network partitions, especially in high-dimensional settings.
- Derives conditional posterior distributions for group assignments using predictive distributions based on sufficient statistics of inter-group edge counts.
- Incorporates likelihoods for both directed and undirected networks through separate parametric forms of the edge probability model $p_\theta(\cdot)$, conditioned on group memberships.
- Computes Metropolis-Hastings acceptance ratios using predictive probabilities and prior densities over partition configurations, enabling efficient transdimensional moves.
Experimental results
Research questions
- RQ1How can we jointly estimate community structures across multiple networks while allowing for structural differences and similarities?
- RQ2What is an effective way to compare community structures across networks in a statistically principled manner?
- RQ3How can we model multiple network types (binary, count, directed, undirected) within a single coherent framework?
- RQ4What MCMC moves are most effective for exploring complex, high-dimensional partition spaces in network blockmodels?
- RQ5How can we ensure posterior inference is robust and well-mixing when the number of communities is unknown and potentially large?
Key findings
- The proposed model successfully identifies shared and distinct community structures across multiple networks, as demonstrated in both simulation studies and real-world data.
- The inclusion of split-merge MCMC moves significantly improves mixing and convergence in the posterior distribution over partitions, outperforming standard Gibbs sampling in high-dimensional settings.
- The model achieves accurate community detection even when networks vary in density, directionality, and data type, showing robustness to model mis-specification.
- The use of a nonparametric prior allows the number of communities to be inferred from data without requiring pre-specification, reducing model selection bias.
- Posterior predictive checks confirm that the model captures realistic network features such as clustering and degree heterogeneity.
- The method enables meaningful aggregation of information across similar networks, improving estimation accuracy compared to independent network analysis.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.