Skip to main content
QUICK REVIEW

[论文解读] Regularized spectral methods for clustering signed networks

Mihai Cucuringu, Apoorv Vikram Singh|arXiv (Cornell University)|Nov 3, 2020
Complex Network Analysis Techniques参考文献 47被引用 13
一句话总结

本文提出用于聚类符号网络的正则化谱方法,将 SPONGE 和符号拉普拉斯算法扩展至具有理论保证的稀疏图。该研究首次对符号随机块模型下的正则化谱聚类提供理论分析,表明在密集和稀疏两种情形下均表现稳健,且在不平衡和稀疏设置下聚类准确率有所提升。

ABSTRACT

We study the problem of $k$-way clustering in signed graphs. Considerable attention in recent years has been devoted to analyzing and modeling signed graphs, where the affinity measure between nodes takes either positive or negative values. Recently, Cucuringu et al. [CDGT 2019] proposed a spectral method, namely SPONGE (Signed Positive over Negative Generalized Eigenproblem), which casts the clustering task as a generalized eigenvalue problem optimizing a suitably defined objective function. This approach is motivated by social balance theory, where the clustering task aims to decompose a given network into disjoint groups, such that individuals within the same group are connected by as many positive edges as possible, while individuals from different groups are mainly connected by negative edges. Through extensive numerical simulations, SPONGE was shown to achieve state-of-the-art empirical performance. On the theoretical front, [CDGT 2019] analyzed SPONGE and the popular Signed Laplacian method under the setting of a Signed Stochastic Block Model (SSBM), for $k=2$ equal-sized clusters, in the regime where the graph is moderately dense. In this work, we build on the results in [CDGT 2019] on two fronts for the normalized versions of SPONGE and the Signed Laplacian. Firstly, for both algorithms, we extend the theoretical analysis in [CDGT 2019] to the general setting of $k \geq 2$ unequal-sized clusters in the moderately dense regime. Secondly, we introduce regularized versions of both methods to handle sparse graphs -- a regime where standard spectral methods underperform -- and provide theoretical guarantees under the same SSBM model. To the best of our knowledge, regularized spectral methods have so far not been considered in the setting of clustering signed graphs. We complement our theoretical results with an extensive set of numerical experiments on synthetic data.

研究动机与目标

  • 将 SPONGE 和符号拉普拉斯算法的理论分析从 k=2 的等大小簇扩展至中等密度符号图中的一般 k ≥ 2 个不等大小簇。
  • 为 SPONGE 和符号拉普拉斯算法开发正则化版本,以在标准谱方法失效的稀疏图情形下提升性能。
  • 在符号随机块模型(SSBM)下为密集和稀疏情形提供理论保证。
  • 提出一种数据驱动的正则化参数选择框架,增强真实世界稀疏符号网络中聚类的鲁棒性。

提出的方法

  • 通过引入正负正则化参数(γ⁺, γ⁻),提出 SPONGE 和对称符号拉普拉斯的正则化版本,以在稀疏图中稳定谱分解。
  • 在 SPONGE sym 中采用广义特征值问题公式,以优化与社会平衡理论一致的符号目标函数。
  • 应用矩阵扰动理论和集中不等式,以边界谱间隙和特征子空间偏离真实聚类结构的程度。
  • 通过在正则化参数(γ⁺, γ⁻)上进行网格搜索来调整性能,并在合成 SSBM 数据上进行经验验证。
  • 提出 SPONGE 算法的归一化版本(SPONGE sym),以提升在不同簇大小下的稳定性和可扩展性。
  • 利用符号随机块模型(SSBM)作为基础生成模型,推导出误聚类率的理论边界。

实验结果

研究问题

  • RQ1SPONGE sym 和符号拉普拉斯算法能否在中等密度情形下,理论扩展至 k ≥ 2 个不等大小簇?
  • RQ2如何使符号网络的谱聚类在稀疏情形下具有鲁棒性,以应对标准方法失效的情况?
  • RQ3正则化参数(γ⁺, γ⁻)对稀疏符号图中聚类性能有何影响?
  • RQ4能否在符号随机块模型下为正则化谱聚类建立理论保证?
  • RQ5正则化算法在不同簇不平衡程度和稀疏水平下的性能特征如何比较?

主要发现

  • 正则化 SPONGE sym 和符号拉普拉斯算法在稀疏图情形下显著提升了聚类准确率,尤其在边密度较低时(p ≤ 0.003)表现突出。
  • 对于 k ≥ 3 且簇大小不平衡的情形,SPONGE sym 在调整兰德指数方面优于所有其他方法,包括未正则化的版本。
  • 正则化算法对簇不平衡具有鲁棒性,即使在高比例(ρ > 1)下性能下降也极小。
  • 参数空间的热力图揭示了正则化符号拉普拉斯算法的显著性能区域,表明由于正负子图密度不同,γ⁺ 和 γ⁻ 的影响具有非对称性。
  • 在 SSBM 下建立了特征子空间偏离和误聚类率的理论边界,证实了在密集和稀疏情形下的收敛性。
  • 对正则化参数(γ⁺, γ⁻)进行网格搜索可带来一致的性能提升,SPONGE sym sparse 展现出与未正则化 SPONGE sym 类似的梯度式性能趋势。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。