Skip to main content
QUICK REVIEW

[论文解读] Degree-degree correlations in random graphs with heavy-tailed degrees

Nelly Litvak, Remco van der Hofstad|University of Twente Research Information|Feb 14, 2012
Hermeneutics and Narrative Identity被引用 4
一句话总结

本文表明,在度数具有重尾分布的无标度网络中,度-度依赖关系的皮尔逊相关系数存在根本性缺陷:即使在网络规模极大时,其值仍可能收敛到非负极限或随机波动,即便在强负相关网络中也是如此。作为替代方案,作者主张采用斯皮尔曼等级相关系数等秩相关度量,这类方法能一致收敛到有意义的极限,并可靠地捕捉不同网络规模下的正负依赖关系。

ABSTRACT

Mixing patterns in large self-organizing networks, such as the Internet, the World Wide Web, social and biological networks are often characterized by degree-degree {dependencies} between neighbouring nodes. One of the problems with the commonly used Pearson's correlation coefficient (termed as the assortativity coefficient) is that {in disassortative networks its magnitude decreases} with the network size. This makes it impossible to compare mixing patterns, for example, in two web crawls of different size. We start with a simple model of two heavy-tailed highly correlated random variable $X$ and $Y$, and show that the sample correlation coefficient converges in distribution either to a proper random variable on $[-1,1]$, or to zero, and if $X,Y\ge 0$ then the limit is non-negative. We next show that it is non-negative in the large graph limit when the degree distribution has an infinite third moment. We consider the alternative degree-degree dependency measure, based on the Spearman's rho, and prove that it converges to an appropriate limit under very general conditions. We verify that these conditions hold in common network models, such as configuration model and Preferential Attachment model. We conclude that rank correlations provide a suitable and informative method for uncovering network mixing patterns.

研究动机与目标

  • 识别在具有重尾度分布的无标度网络中使用皮尔逊相关系数测量度-度依赖关系的根本缺陷。
  • 证明即使在存在强烈负度-度依赖关系的网络中,皮尔逊相关系数仍可能收敛到非负极限或随机变量。
  • 提出基于秩的依赖度量(尤其是斯皮尔曼等级相关系数)作为更可靠的替代方法,该方法对有限矩的数量不敏感,并在大网络极限下具有有意义的收敛性。
  • 建立理论条件,以确保在典型网络模型(如配置模型和优先连接模型)中,斯皮尔曼等级相关系数收敛到明确定义的极限。
  • 为在具有幂律度分布的复杂网络背景下使用基于秩的依赖度量而非皮尔逊相关系数提供理论依据。

提出的方法

  • 从理论上研究当两个具有无限三阶矩的重尾、高度相关随机变量时,样本皮尔逊相关系数的渐近行为。
  • 将重尾随机变量的相关结果应用于随机图中的度-度依赖关系,证明当度数非负且具有无限方差时,皮尔逊相关系数在渐近下非负。
  • 构造明确的实例,表明在度-度负相关的网络中,皮尔逊相关系数可能收敛到零,或在分布上收敛到一个非退化的随机变量,尽管存在强烈的负依赖。
  • 提出斯皮尔曼等级相关系数作为基于节点度数秩变换的替代依赖度量,以避免对极端值的敏感性。
  • 证明在度分布的一般条件下,斯皮尔曼等级相关系数收敛到一个明确定义的极限,包括在配置模型和优先连接模型中所满足的条件。
  • 运用极值理论和拷贝理论的工具,特别是“角测度”概念,来解释尾部依赖性,并与秩相关结果进行比较。

实验结果

研究问题

  • RQ1为何皮尔逊相关系数在具有重尾度的大型无标度网络中无法可靠地测量度-度依赖关系?
  • RQ2在何种条件下,皮尔逊相关系数会收敛到非负极限,即使真实依赖关系为强烈负相关?
  • RQ3基于秩的相关度量(如斯皮尔曼等级相关系数)是否能在不同网络规模下一致捕捉正负度-度依赖关系?
  • RQ4在具有幂律度分布的随机图模型中,何种理论条件可确保斯皮尔曼等级相关系数收敛到有意义的极限?
  • RQ5在捕捉复杂网络中的尾部依赖性和结构混合模式方面,基于秩的度量与皮尔逊相关系数相比有何差异?

主要发现

  • 在具有无限三阶矩的无标度网络中,度-度依赖关系的皮尔逊相关系数在分布上收敛到非负随机变量或零,即使真实依赖关系为强烈负相关。
  • 在具有重尾度的网络中,随着网络规模增大,皮尔逊相关系数可能持续无限波动,导致其在不同规模网络间比较混合模式时不可靠。
  • 在度分布的一般条件下,斯皮尔曼等级相关系数一致收敛到一个明确定义的极限,包括在配置模型和优先连接模型中。
  • 使用秩相关度量可避免皮尔逊相关系数的病态行为,并提供一种对度分布有限矩数量不敏感的稳定、信息丰富的依赖度量。
  • 本文为为何在具有重尾度的网络中,基于秩的相关度量比皮尔逊相关系数更合适提供了理论依据,尤其是在研究负度相关结构时。
  • 结果表明,基于秩的相关度量对极端值具有鲁棒性,即使在网络规模趋于无穷的极限下,也能可靠检测正负度-度依赖关系。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。