[论文解读] Using Maximum Entry-Wise Deviation to Test the Goodness of Fit for Stochastic Block Models
本文提出了一种基于中心化并标准化邻接矩阵的最大逐元素偏差的拟合优度检验,用于随机块模型。在原假设下,该检验渐近服从极值分布(Gumbel分布),允许社区数 $k$ 随节点数 $n$ 线性增长(最多至 $k = o(n / \log^2 n)$),显著放宽了以往的限制,并在一类替代假设下保持渐近功效。
The stochastic block model is widely used for detecting community structures in network data. How to test the goodness-of-fit of the model is one of the fundamental problems and has gained growing interests in recent years. In this article, we propose a novel goodness-of-fit test based on the maximum entry of the centered and re-scaled adjacency matrix for the stochastic block model. One noticeable advantage of the proposed test is that the number of communities can be allowed to grow linearly with the number of nodes ignoring a logarithmic factor. We prove that the null distribution of the test statistic converges in distribution to a Gumbel distribution, and we show that both the number of communities and the membership vector can be tested via the proposed method. Further, we show that the proposed test has asymptotic power guarantee against a class of alternatives. We also demonstrate that the proposed method can be extended to the degree-corrected stochastic block model. Both simulation studies and real-world data examples indicate that the proposed method works well.
研究动机与目标
- 为解决在社区数随网络规模增长时,随机块模型中模型拟合检验的根本性挑战。
- 开发一种在社区数 $k$ 随节点数 $n$ 增加时仍保持有效性与强大功效的检验方法,突破以往的限制。
- 将该方法扩展至度校正的随机块模型,以考虑节点度的异质性。
- 在最小假设下,提供检验统计量渐近零分布与功效的理论保证。
- 提供一种计算高效的替代方案,相较于蒙特卡洛或似然方法,具备理论支持。
提出的方法
- 基于中心化并标准化邻接矩阵 $A_{ij} - B_{\sigma(i)\sigma(j)}$ 的最大绝对逐元素偏差,定义检验统计量,并除以其标准差。
- 证明在原假设下,当 $k = o(n / \log^2 n)$ 时,检验统计量依分布收敛至Gumbel分布,从而允许 $k$ 几乎线性地随 $n$ 增长。
- 提出一种增强的检验统计量,可在不改变渐近零分布的前提下提升检验功效。
- 利用极值理论与中偏差不等式推导极限分布,利用弱依赖但存在依赖关系的条目最大值。
- 通过引入估计的度参数 $\widehat{\omega}_i$,将框架扩展至度校正的随机块模型,但其渐近零分布仍为开放问题。
- 在有限样本中应用自助法(bootstrapping)计算 p 值,尤其适用于真实数据应用。
实验结果
研究问题
- RQ1能否为随机块模型开发一种拟合优度检验,使其在社区数随节点数增长时仍保持有效性?
- RQ2在随机块模型下,邻接矩阵中最大逐元素偏差的渐近零分布为何种分布?
- RQ3所提出的检验是否对一类局部替代假设保持渐近功效?
- RQ4该方法能否扩展至具有未知度参数的更灵活的度校正随机块模型?
- RQ5与蒙特卡洛或似然比检验等现有方法相比,该检验在实际中表现如何?
主要发现
- 当 $k = o(n / \log^2 n)$ 时,基于最大逐元素偏差的检验统计量在原假设下依分布收敛至Gumbel分布,允许 $k$ 几乎线性地随 $n$ 增长。
- 所提出的检验对一类局部替代假设具有渐近功效,确保当偏差可检测时能够识别出模型拟合不佳。
- 增强的检验统计量在不改变渐近零分布的前提下,提升了经验功效。
- 在政治博客网络中,标准随机块模型下,$k=10$ 被拒绝($T^{+}_{n,\text{boot}} = 23.59 > 3.41$),而度校正模型下 $k=2$ 未被拒绝($T^{+}_{n2,\text{boot}} = 2.06 < 3.41$),与先前发现一致。
- 在理论灵活性方面,该方法优于现有方法,尤其在 $k$ 随 $n$ 增长的高维设置中表现更优。
- 由于估计的度参数引入了复杂依赖关系,度校正模型的渐近零分布仍未知,凸显了该领域的一个关键开放问题。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。