Skip to main content
QUICK REVIEW

[论文解读] Decoding the structure of the WWW: facts versus sampling biases

M. Ángeles Serrano, Ana Gabriela Maguitman|ArXiv.org|Nov 8, 2005
Web visibility and informetrics参考文献 32被引用 6
一句话总结

本文通过分析四个大规模、来源各异的WWW抓取数据集,研究了网络图测量中的采样偏差。研究识别出互惠链接(页面之间的相互超链接)作为关键的拓扑特征,揭示了结构相关性,并将真实的网络组织结构与爬虫引入的伪影区分开来,证明互惠链接的统计特性承载了网络图中绝大部分有意义的拓扑信息。

ABSTRACT

The understanding of the immense and intricate topological structure of the World Wide Web (WWW) is a major scientific and technological challenge. This has been tackled recently by characterizing the properties of its representative graphs in which vertices and directed edges are identified with web-pages and hyperlinks, respectively. Data gathered in large scale crawls have been analyzed by several groups resulting in a general picture of the WWW that encompasses many of the complex properties typical of rapidly evolving networks. In this paper, we report a detailed statistical analysis of the topological properties of four different WWW graphs obtained with different crawlers. We find that, despite the very large size of the samples, the statistical measures characterizing these graphs differ quantitatively, and in some cases qualitatively, depending on the domain analyzed and the crawl used for gathering the data. This spurs the issue of the presence of sampling biases and structural differences of Web crawls that might induce properties not representative of the actual global underlying graph. In order to provide a more accurate characterization of the Web graph and identify observables which are clearly discriminating with respect to the sampling process, we study the behavior of degree-degree correlation functions and the statistics of reciprocal connections. The latter appears to enclose the relevant correlations of the WWW graph and carry most of the topological information of theWeb. The analysis of this quantity is also of major interest in relation to the navigability and searchability of the Web.

研究动机与目标

  • 评估不同爬取策略对网络图测量拓扑特性的影响。
  • 识别在多种数据样本中保持稳健的结构特征,以及受采样偏差影响的伪影特征。
  • 确定某些统计度量(尤其是互惠链接结构)是否比标准度分布更能可靠地表征真实网络拓扑。
  • 评估互惠链接在塑造网络可导航性和可搜索性方面的作用。

提出的方法

  • 对四个来自不同领域和时间段的独立抓取的大规模WWW图进行比较性统计分析。
  • 检查入度和出度分布、度-度相关性以及聚类系数。
  • 通过将互惠链接子图视为无向图,重点关注其局部连通性,利用聚类系数 $ c_j $ 和平均最近邻度 $ \overline{q}_{r,nn}(q_r) $ 进行分析。
  • 使用聚类系数 $ \overline{c}(q_r) $ 评估互惠子图中邻域的互联程度。
  • 分析度依赖的聚类系数和最近邻度,以检测如团或星型结构等结构性模式。
  • 通过多个抓取数据集的交叉验证,区分持久的结构性特征与采样伪影。

实验结果

研究问题

  • RQ1不同网络爬虫在多大程度上产生定性和定量上不同的网络图拓扑测量结果?
  • RQ2哪些拓扑可观测量在多种抓取数据集中保持一致,哪些对采样偏差敏感?
  • RQ3互惠链接在整体网络图结构和相关性模式中起到何种作用?
  • RQ4互惠链接子图能否作为理解真实网络底层拓扑结构的可靠代理?
  • RQ5互惠链接在决定网络可导航性和可搜索性方面发挥什么作用?

主要发现

  • 互惠链接子图表现出高且恒定的聚类系数 $ \overline{c}(q_r) $,表明存在高度互联的社区核心结构和星型配置,且与抓取来源无关。
  • 在所有四个数据集中,互惠子图中的平均最近邻度 $ \overline{q}_{r,nn}(q_r) $ 和聚类系数 $ \overline{c}(q_r) $ 均表现出一致的模式,表明这是网络的真实结构性特征。
  • 非互惠的度-度相关性在不同抓取数据中差异显著,表明标准度相关度量对采样偏差高度敏感。
  • 互惠链接统计是最具信息量的拓扑可观测量,承载了网络图中大部分相关的结构信息。
  • 密集且高度互联的互惠子图(特别是团和星型枢纽)的存在,表明其在网页导航和抗链接失效方面具有功能性作用。
  • 尽管数据规模庞大,网络图的观测拓扑特性仍显著受爬虫设计影响,这削弱了先前研究中关于普遍性的主张。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。