Skip to main content
QUICK REVIEW

[论文解读] Determining the Number of Communities in Degree-corrected Stochastic Block Models

Shujie Ma, Liangjun Su|arXiv (Cornell University)|Sep 4, 2018
Complex Network Analysis Techniques参考文献 32被引用 19
一句话总结

本文提出一种伪似然比统计量,结合谱聚类与二分分割法,以一致地估计度修正随机块模型中的社区数。在平均度至少以 log(n) 的速度增长的温和条件下,该方法确保了一致性,并推导出过拟合与欠拟合情形下的极限分布。

ABSTRACT

We propose to estimate the number of communities in degree-corrected stochastic block models based on a pseudo likelihood ratio statistic. To this end, we introduce a method that combines spectral clustering with binary segmentation. This approach guarantees an upper bound for the pseudo likelihood ratio statistic when the model is over-fitted. We also derive its limiting distribution when the model is under-fitted. Based on these properties, we establish the consistency of our estimator for the true number of communities. Developing these theoretical properties require a mild condition on the average degrees -- growing at a rate no slower than log(n), where n is the number of nodes. Our proposed method is further illustrated by simulation studies and analysis of real-world networks. The numerical results show that our approach has satisfactory performance when the network is semi-dense.

研究动机与目标

  • 为解决在实际中通常未知真实社区数的度修正随机块模型(DCSBM)中估计真实社区数这一关键挑战。
  • 开发一种利用完整数据结构而非依赖特征值等部分统计量的模型选择方法。
  • 在平均度增长率不低于 log(n) 的最小条件下,建立估计量的理论一致性。
  • 为 DCSBM 设置提供一种计算可行且理论基础坚实的替代方案,以替代迭代或随机化方法(如交叉验证)。
  • 推导在过拟合与欠拟合条件下伪似然比统计量的极限分布,以支持模型选择。

提出的方法

  • 基于谱聚类提出一种伪似然比统计量,用于评估不同社区数下的模型拟合程度。
  • 引入一种二分分割程序,从较高的上界开始,迭代测试候选的社区数。
  • 对每个候选 K 值,使用谱聚类估计社区成员身份及潜在参数(如社区均值与度数)。
  • 在模型过拟合时,建立伪似然比统计量的上界,以确保选择过程的稳定性。
  • 推导在模型欠拟合条件下的伪似然比统计量极限分布,以支持渐近推断。
  • 在谱嵌入上应用正则化技术与 K-均值聚类,以提高社区分配的准确性。

实验结果

研究问题

  • RQ1在现实条件下,能否使用伪似然比统计量一致地估计 DCSBM 中的社区数?
  • RQ2如何结合谱聚类与二分分割法,以确保社区检测模型选择的理论一致性?
  • RQ3当模型欠拟合时,伪似然比统计量的极限分布是什么?
  • RQ4为保证估计量的一致性,对平均度的最小条件是什么?
  • RQ5与现有方法(如交叉验证或基于特征值的方法)相比,该方法在理论支持与计算稳定性方面表现如何?

主要发现

  • 在平均度至少以 log(n) 的速度增长的条件下,所提出的估计量对真实社区数具有一致性。
  • 当模型过拟合时,伪似然比统计量有上界,确保方法不会高估 K。
  • 在欠拟合情形下,推导出伪似然比统计量的极限分布,支持渐近推断。
  • 即使社区数未知且需从数据中选择,该方法仍能实现一致的社区检测。
  • 模拟研究显示,在半密集网络中表现令人满意,K 的估计稳定且准确。
  • 理论结果得到附录中严格证明的支持,包括关于聚类误差与参数估计的技术引理。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。