[论文解读] Learning Social Circles in Ego Networks based on Multi-View Social Graphs
本文提出了一种改进的多视角谱聚类方法,用于利用来自 Twitter 的六种结构视角——用户关系、互动行为和内容——检测自我中心网络中的社交圈。结果表明,视角特异性的信息迁移可提升在稀疏或潜在不完整网络上的鲁棒性,优于标准的多视角聚类方法,且在自动与人工评估中均得到实证验证。
In social network analysis, automatic social circle detection in ego-networks is becoming a fundamental and important task, with many potential applications such as user privacy protection or interest group recommendation. So far, most studies have focused on addressing two questions, namely, how to detect overlapping circles and how to detect circles using a combination of network structure and network node attributes. This paper asks an orthogonal research question, that is, how to detect circles based on network structures that are (usually) described by multiple views? Our investigation begins with crawling ego-networks from Twitter and employing classic techniques to model their structures by six views, including user relationships, user interactions and user content. We then apply both standard and our modified multi-view spectral clustering techniques to detect social circles in these ego-networks. Based on extensive automatic and manual experimental evaluations, we deliver two major findings: first, multi-view clustering techniques perform better than common single-view clustering techniques, which only use one view or naively integrate all views for detection, second, the standard multi-view clustering technique is less robust than our modified technique, which selectively transfers information across views based on an assumption that sparse network structures are (potentially) incomplete. In particular, the second finding makes us believe a direct application of standard clustering on potentially incomplete networks may yield biased results. We lightly examine this issue in theory, where we derive an upper bound for such bias by integrating theories of spectral clustering and matrix perturbation, and discuss how it may be affected by several network characteristics.
研究动机与目标
- 解决在使用多个可能不完整的视角表示网络结构时,检测自我中心网络中社交圈的挑战。
- 通过选择性地在视角之间传递信息而非简单地组合所有视角,提升在稀疏自我中心网络中的聚类鲁棒性。
- 探究标准多视角聚类方法在应用于不完整或稀疏网络数据时是否存在偏差。
提出的方法
- 从 Twitter 抓取自我中心网络,并使用六种不同的视角(用户关系、互动行为和内容)对其结构进行建模。
- 使用通过平均各视角特异性相似度矩阵构建的共识图,应用标准多视角谱聚类。
- 提出一种改进的多视角谱聚类方法,基于视角的可靠性与稀疏性,选择性地在视角之间传递信息。
- 采用基于矩阵扰动的理论分析,界定不完整网络结构在聚类中引入的偏差。
- 通过自动指标与在多样化自我中心网络中对检测到的社交圈进行人工标注,评估性能。
- 将谱聚类与视角特异性权重调整相结合,以增强对稀疏或不完整视角的鲁棒性。
实验结果
研究问题
- RQ1多视角聚类在检测自我中心网络中重叠社交圈时,与单视角聚类相比表现如何?
- RQ2与简单融合所有视角的方法相比,选择性地在视角之间传递信息是否能提升聚类的鲁棒性?
- RQ3网络稀疏性与不完整性对多视角设置中聚类偏差的理论影响是什么?
- RQ4不同网络特征如何影响多视角聚类中偏差的上界?
主要发现
- 多视角聚类技术在检测自我中心网络中的社交圈方面,显著优于单视角方法。
- 所提出的改进多视角聚类方法在稀疏且可能不完整的网络结构上,相比标准多视角聚类展现出更好的鲁棒性。
- 理论分析表明,标准多视角聚类在应用于不完整网络时容易产生偏差,该分析基于谱聚类与矩阵扰动理论。
- 理论偏差上界随网络稀疏性增加而上升,随视角一致性提高而降低,表明视角可靠性影响偏差的大小。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。