[Paper Review] Hypothesis Testing for Equality of Latent Positions in Random Graphs
This paper proposes hypothesis tests for equality of latent positions in generalized random dot product graphs using Mahalanobis distances from spectral embeddings of the adjacency or normalized Laplacian matrices. Under mild regularity conditions, the test statistics asymptotically follow chi-square distributions under both the null and local alternative, enabling model selection between stochastic block models, degree-corrected variants, and Erdős–Rényi models with theoretical non-centrality parameters derived for power analysis.
We consider the hypothesis testing problem that two vertices $i$ and $j$ of a generalized random dot product graph have the same latent positions, possibly up to scaling. Special cases of this hypothesis test include testing whether two vertices in a stochastic block model or degree-corrected stochastic block model graph have the same block membership vectors, or testing whether two vertices in a popularity adjusted block model have the same community assignment. We propose several test statistics based on the empirical Mahalanobis distances between the $i$th and $j$th rows of either the adjacency or the normalized Laplacian spectral embedding of the graph. We show that, under mild conditions, these test statistics have limiting chi-square distributions under both the null and local alternative hypothesis, and we derived explicit expressions for the non-centrality parameters under the local alternative. Using these limit results, we address the model selection problems including choosing between the standard stochastic block model and its degree-corrected variant, and choosing between the ER model and stochastic block model. The effectiveness of our proposed tests are illustrated via both simulation studies and real data applications.
Motivation & Objective
- To develop statistical hypothesis tests for whether two vertices in a random graph have equal latent positions, up to scaling.
- To address model selection challenges between stochastic block models (SBM), degree-corrected SBMs (DC-SBM), and Erdős–Rényi (ER) models.
- To provide asymptotically valid inference for vertex-level equality in latent positions under the generalized random dot product graph (GRDPG) framework.
- To enable applications in vertex nomination and role discovery by testing structural similarity between nodes.
- To derive non-centrality parameters under local alternatives for power analysis and empirical validation.
Proposed method
- Proposes test statistics based on empirical Mahalanobis distances between the i-th and j-th rows of spectral embeddings from the adjacency matrix or normalized Laplacian matrix.
- Uses spectral embedding of the graph matrix to obtain consistent estimates of latent positions under the GRDPG model.
- Derives asymptotic chi-square distributions for test statistics under both the null hypothesis and local alternative, with explicit non-centrality parameters.
- Applies the limiting distributions to conduct hypothesis tests on vertex-level equality of latent positions, including up to scaling.
- Employs Monte Carlo simulations with 500 replicates to evaluate empirical size and power across varying sparsity levels.
- Validates theoretical results using real data on political blogs, demonstrating close alignment between empirical and theoretical power.
Experimental results
Research questions
- RQ1Can we test whether two vertices in a generalized random dot product graph have the same latent positions, possibly up to scaling, using spectral embeddings?
- RQ2What is the asymptotic distribution of Mahalanobis distance-based test statistics for vertex pair equality under the null and local alternative hypotheses?
- RQ3How can we use these test statistics to select between competing network models such as SBM, DC-SBM, and ER models?
- RQ4What are the explicit non-centrality parameters for the non-central chi-square distribution under the local alternative, and how do they affect test power?
- RQ5How well do the asymptotic approximations perform in finite samples, especially under varying sparsity levels?
Key findings
- The test statistics $ T_{ ext{out}} $, $ T_{ ext{in}} $, and $ T_{ ext{both}} $ asymptotically follow $ \chi^2_3 $, $ \chi^2_3 $, and $ \chi^2_6 $ distributions under the null hypothesis, respectively, with empirical histograms closely matching theoretical chi-square densities.
- Empirical size estimates for $ T_{ ext{out}} $ and $ T_{ ext{both}} $ are close to the nominal 0.05 level across all sparsity levels, confirming correct size control.
- Empirical power for $ T_{ ext{in}} $ and $ T_{ ext{both}} $ increases with sparsity, reaching 0.956 and 0.844 at $ \rho = 1.0 $, respectively, closely matching theoretical power predictions.
- Theoretical non-centrality parameters for the local alternative are accurately estimated and correlate strongly with observed power, validating the asymptotic theory.
- In the political blogs dataset, the proposed tests successfully identify structural similarities between vertices, supporting their utility in real-world applications like vertex nomination.
- The method enables effective model selection, distinguishing between SBM, DC-SBM, and ER models based on vertex-level latent position equality tests.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.