[论文解读] Stochastic Block Models for Multiplex networks: an application to networks of researchers
本文提出一种多层随机块模型(MSBM),用于联合分析研究人员之间多种类型的关系,例如直接建议关系与机构隶属关系。通过变分期望-最大化算法,该模型估计块分配及块间连接概率,在具有强关系类型相互依赖性的法国癌症研究网络中展现出一致性,并识别出有意义的聚类。
Modeling relations between individuals is a classical question in social sciences and clustering individuals according to the observed patterns of interactions allows to uncover a latent structure in the data. Stochastic block model (SBM) is a popular approach for grouping the individuals with respect to their social comportment. When several relationships of various types can occur jointly between the individuals, the data are represented by multiplex networks where more than one edge can exist between the nodes. In this paper, we extend the SBM to multiplex networks in order to obtain a clustering based on more than one kind of relationship. We propose to estimate the parameters --such as the marginal probabilities of assignment to groups (blocks) and the matrix of probabilities of connections between groups-- through a variational Expectation-Maximization procedure. Consistency of the estimates as well as statistical properties of the model are obtained. The number of groups is chosen thanks to the Integrated Completed Likelihood criteria, a penalized likelihood criterion. Multiplex Stochastic Block Model arises in many situations but our applied example is motivated by a network of French cancer researchers. The two possible links (edges) between researchers are a direct connection or a connection through their labs. Our results show strong interactions between these two kinds of connections and the groups that are obtained are discussed to emphasize the common features of researchers grouped together.
研究动机与目标
- 开发一种统计模型,联合分析网络中多种社会关系类型,例如直接联系与机构隶属关系。
- 将随机块模型(SBM)扩展至多层网络,其中节点之间存在多种边类型共存。
- 使用变分EM过程估计块成员概率和组间连接概率等参数。
- 采用集成完整似然(ICL)准则选择最优块数。
- 将该模型应用于法国癌症研究人员的真实网络,以揭示建议关系与机构隶属关系中潜在的结构模式。
提出的方法
- 通过将K种独立边类型(例如建议关系与实验室隶属关系)建模为K维伯努利过程,将单层SBM扩展至多层网络。
- 通过类别潜变量τ对每个节点的块成员身份进行建模,其先验概率为αq(对应块q)。
- 定义联合概率模型,其中K类边的存在与否由块对特定的概率πql(w)决定,其中w ∈ {0,1}K表示边存在/不存在的模式。
- 使用变分EM算法最大化观测似然的下界,E步估计潜变量块成员τ,M步更新参数α和π。
- 对τ优化实现定点算法,对α和π在变分似然下采用闭式更新。
- 采用集成完整似然(ICL)准则选择最优块数,对模型复杂度施加惩罚。
实验结果
研究问题
- RQ1如何将随机块模型扩展以在单一统计框架中处理多个相互依赖的网络层?
- RQ2将个体层面的直接联系与机构层面的隶属关系相结合,对识别有意义的社会聚类有何影响?
- RQ3在具有多种边类型和潜变量块结构的多层SBM中,如何实现参数估计的一致性?
- RQ4不同类型关系之间的相互依赖性(例如建议关系与实验室关系)在塑造网络结构中起什么作用?
- RQ5在多层网络设置中,如何可靠地选择潜变量块的数量?
主要发现
- 变分EM算法能一致地估计块成员身份与连接概率,在正则性条件下已证明其理论一致性。
- 该模型在法国癌症研究网络中识别出直接建议关系与机构隶属关系之间存在显著的相互依赖性。
- 集成完整似然(ICL)准则成功选出了最优块数,有效平衡了模型拟合度与复杂度。
- 估计的块结构揭示了具有相似合作与机构模式的研究人员有意义的分组。
- 该模型通过捕捉中尺度聚类,检测到超越规则网络结构(如无标度或小世界特性)的非平凡拓扑特征。
- 在真实数据上的应用表明,同一块内的研究人员在个体网络与机构网络中往往具有相似的位置,暗示在资源获取方面存在结构优势。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。