Skip to main content
QUICK REVIEW

[论文解读] Completely random measures for modeling power laws in sparse graphs

Diana Cai, Tamara Broderick|arXiv (Cornell University)|Mar 22, 2016
Complex Network Analysis Techniques参考文献 23被引用 5
一句话总结

本文提出了一种基于完全随机测度(completely random measures)的新型生成模型,特别采用三参数β过程来捕捉现实网络中观察到的多种幂律行为。该模型实现了稀疏性,并在边数缩放、度为1的顶点数以及度分布方面分别表现出I型、IIa型和IIIa型幂律,模拟结果验证了其幂律斜率分别为1.2、1.1和-1.6。

ABSTRACT

Network data appear in a number of applications, such as online social networks and biological networks, and there is growing interest in both developing models for networks as well as studying the properties of such data. Since individual network datasets continue to grow in size, it is necessary to develop models that accurately represent the real-life scaling properties of networks. One behavior of interest is having a power law in the degree distribution. However, other types of power laws that have been observed empirically and considered for applications such as clustering and feature allocation models have not been studied as frequently in models for graph data. In this paper, we enumerate desirable asymptotic behavior that may be of interest for modeling graph data, including sparsity and several types of power laws. We outline a general framework for graph generative models using completely random measures; by contrast to the pioneering work of Caron and Fox (2015), we consider instantiating more of the existing atoms of the random measure as the dataset size increases rather than adding new atoms to the measure. We see that these two models can be complementary; they respectively yield interpretations as (1) time passing among existing members of a network and (2) new individuals joining a network. We detail a particular instance of this framework and show simulated results that suggest this model exhibits some desirable asymptotic power-law behavior.

研究动机与目标

  • 开发一种稀疏图的生成模型,以捕捉现实网络中观察到的真实缩放行为。
  • 解决现有模型因Aldous–Hoover定理假设而导致生成稠密图的局限性。
  • 通过整合度分布、聚类和特征分配等多样化的幂律行为,将网络模型工具箱的适用范围从密度扩展至更广泛的行为。
  • 提供一种灵活的框架,利用完全随机测度(CRM),可使用不同的基测度实例化,包括三参数β过程。
  • 通过实验和理论方法,探究该模型在稀疏图生成中是否表现出理想的渐近幂律缩放特性。

提出的方法

  • 该框架使用完全随机测度(CRMs)通过从CRM中抽取的权重进行独立伯努利试验,来定义图中边的概率。
  • 该模型指定三参数β过程作为底层CRM,采用杆破除构造方法生成原子及其权重。
  • 为使计算可行,将原子数量截断为5000,参数设置为γ=3,θ=1,α=0.1。
  • 通过采样N个顶点并基于CRM中权重的乘积分配边概率来生成图。
  • 随着N增加,模型允许现有原子被实例化,从而建模现有节点之间的演化,而非新增节点。
  • 在N=50至2000个顶点的范围内进行模拟,以分析缩放行为和幂律拟合。

实验结果

研究问题

  • RQ1基于CRM的生成模型能否生成避免Aldous–Hoover定理所暗示的稠密图行为的稀疏图?
  • RQ2该模型是否在顶点数与边数之间表现出I型幂律缩放?
  • RQ3该模型能否在顶点数与度为1的顶点数之间生成IIa型幂律?
  • RQ4度分布是否遵循IIIa型幂律,如现实网络中所见?
  • RQ5该模型能否扩展以捕捉其他类型的幂律(如IIb型和IIIb型),其理论渐近性质如何?

主要发现

  • 该模型生成了稀疏图,表现为顶点数与边数之间呈次二次缩放,斜率为1.2,表明为I型幂律。
  • 度为1的顶点数随总顶点数的增加而按IIa型幂律缩放,在高N区域斜率为1.1。
  • 度分布表现出IIIa型幂律,低度顶点的斜率为-1.6,与现实网络一致。
  • 模拟结果证实,该模型能同时捕捉多种幂律行为,表明其在建模复杂网络结构方面具有潜力。
  • 该模型的行为与现有模型互补:其将增长解释为现有节点之间的演化,而非新节点的加入。
  • 初步结果表明,该模型是未来理论分析和现实稀疏网络实际推断的有前途候选。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。