Skip to main content
QUICK REVIEW

[论文解读] Hierarchical Stochastic Block Model for Community Detection in Multiplex Networks

Arash A. Amini, Marina Silva Paez|arXiv (Cornell University)|Mar 30, 2019
Bayesian Methods and Mixture Models参考文献 41被引用 7
一句话总结

本文提出一种分层随机块模型(HSBM),用于在多层网络中进行社区检测,利用分层狄利克雷过程先验,使不同层的社区结构可变,同时在层间共享信息。该方法可自动选择各层的社区数量,在模拟和真实世界网络(包括FAO贸易网络)中均能检测到有意义且稳定的社区结构,优于单层模型。

ABSTRACT

Multiplex networks have become increasingly more prevalent in many fields, and have emerged as a powerful tool for modeling the complexity of real networks. There is a critical need for developing inference models for multiplex networks that can take into account potential dependencies across different layers, particularly when the aim is community detection. We add to a limited literature by proposing a novel and efficient Bayesian model for community detection in multiplex networks. A key feature of our approach is the ability to model varying communities at different network layers. In contrast, many existing models assume the same communities for all layers. Moreover, our model automatically picks up the necessary number of communities at each layer (as validated by real data examples). This is appealing, since deciding the number of communities is a challenging aspect of community detection, and especially so in the multiplex setting, if one allows the communities to change across layers. Borrowing ideas from hierarchical Bayesian modeling, we use a hierarchical Dirichlet prior to model community labels across layers, allowing dependency in their structure. Given the community labels, a stochastic block model (SBM) is assumed for each layer. We develop an efficient slice sampler for sampling the posterior distribution of the community labels as well as the link probabilities between communities. In doing so, we address some unique challenges posed by coupling the complex likelihood of SBM with the hierarchical nature of the prior on the labels. An extensive empirical validation is performed on simulated and real data, demonstrating the superior performance of the model over single-layer alternatives, as well as the ability to uncover interesting structures in real networks.

研究动机与目标

  • 解决现有模型假设多层网络中所有层社区结构相同的局限性。
  • 开发一种灵活的贝叶斯模型,允许社区结构在各层间变化,同时捕捉结构依赖性。
  • 自动确定每层的社区数量,避免预先指定。
  • 通过分层先验在层间共享信息,提升社区检测的准确性。
  • 为多层SBM设置中复杂耦合似然提供高效的后验推断算法。

提出的方法

  • 使用分层狄利克雷过程(HDP)作为网络各层社区标签的非参数先验。
  • 对每一层分配一个随机块模型(SBM),条件于社区标签。
  • 采用高效的切片抽样法进行社区标签和层间社区连接概率的后验推断。
  • 通过随机划分建模社区结构,实现在数据驱动下的灵活社区检测。
  • 将SBM似然与分层先验耦合,联合估计社区结构和连接概率。
  • 应用Fruchterman–Reingold布局算法,基于各层网络连通性可视化节点位置。

实验结果

研究问题

  • RQ1贝叶斯模型能否在社区结构随层变化的多层网络中有效检测社区?
  • RQ2允许各层具有特定社区结构,相比假设社区共享的模型,能否提升检测准确性?
  • RQ3分层先验在稀疏或复杂多层网络中,能在多大程度上通过跨层共享信息改善估计?
  • RQ4该模型能否在无需预设的情况下自动选择每层的社区数量?
  • RQ5该模型在真实世界多层网络中面对稀疏性和噪声时,表现如何?

主要发现

  • HSBM模型在估计群体内部的平均归一化汉明(ANH)距离显著低于随机配对,样本内组内配对的中位ANH为0.027,而随机配对为0.21。
  • 样本外ANH结果证实了模型的稳健性,即使在层更稀疏的情况下,组内与随机配对分布仍保持一致分离。
  • 识别出反直觉的国家组合,如澳门–卢旺达(ANH = 0.074)和伊拉克–几内亚(ANH = 0.019),表明其在食品产品类别间存在共享贸易模式。
  • 该模型成功揭示了FAO贸易网络中具有地理和经济意义的聚类,包括在多层中频繁出现标签的群体(例如,加拿大在13/20层中属于第6组)。
  • 切片抽样法在SBM似然与分层先验复杂耦合的条件下,仍实现了高效的后验抽样,支持可扩展的推断。
  • 在实证验证中,HSBM优于单层替代模型,展现出在检测稳定且可解释的社区结构方面的卓越性能。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。