Skip to main content
QUICK REVIEW

[论文解读] On the Ubiquity of Web Tracking: Insights from a Billion-Page Web Crawl

Sebastian Schelter, Jérôme Kunegis|arXiv (Cornell University)|Jul 25, 2016
Privacy, Security, and Data Protection参考文献 25被引用 4
一句话总结

本研究基于 CommonCrawl 2012 的十亿页网页抓取数据,对 4100 万个域名中的 1.4 亿个追踪器嵌入进行了分析,揭示了谷歌、脸书和推特在全球范围内的主导地位。政治因素(如新闻自由)解释了这些偏差,尤其是在中国和俄罗斯;尽管隐私敏感型网站的追踪器使用率仍高达 60%,但低于普通网站的水平。

ABSTRACT

We perform a large-scale analysis of third-party trackers on the World Wide Web from more than 3.5 billion web pages of the CommonCrawl 2012 corpus. We extract a dataset containing more than 140 million third-party embeddings in over 41 million domains. To the best of our knowledge, this constitutes the largest web tracking dataset collected so far, and exceeds related studies by more than an order of magnitude in the number of domains and web pages analyzed. We perform a large-scale study of online tracking, on three levels: (1) On a global level, we give a precise figure for the extent of tracking, give insights into the structure of the `online tracking sphere' and analyse which trackers are used by how many websites. (2) On a country-specific level, we analyse which trackers are used by websites in different countries, and identify the countries in which websites choose significantly different trackers than in the rest of the world. (3) We answer the question whether the content of websites influences the choice of trackers they use, leveraging more than 90 thousand categorized domains. In particular, we analyse whether highly privacy-critical websites make different choices of trackers than other websites. Based on the performed analyses, we confirm that trackers are widespread (as expected), and that a small number of trackers dominates the web (Google, Facebook and Twitter). In particular, the three tracking domains with the highest PageRank are all owned by Google. The only exception to this pattern are a few countries such as China and Russia. Our results suggest that this dominance is strongly associated with country-specific political factors such as freedom of the press. We also confirm that websites with highly privacy-critical content are less likely to contain trackers (60% vs 90% for other websites), even though the majority of them still do contain trackers.

研究动机与目标

  • 利用大规模数据集,理解第三方网络追踪的全球规模与结构特性。
  • 调查追踪器使用在各国之间的差异,并探究政治、文化或经济因素是否影响这些模式。
  • 检验涵盖隐私敏感主题(如健康与成瘾)的网站是否比其他网站使用更少的追踪器。
  • 识别主导追踪服务,并评估其在网页中的覆盖范围,尤其关注地缘政治因素的影响。

提出的方法

  • 从 CommonCrawl 2012 语料库中的 35 亿个网页中提取第三方追踪器嵌入,聚焦于付费级别域名(PLDs)。
  • 构建一个双分图追踪网络,将网站(PLDs)与第三方追踪服务关联,支持大规模网络分析。
  • 应用包括 PageRank、度分布和异配性在内的网络分析技术,研究追踪领域的结构特性。
  • 通过共现模式聚类,识别国家特定和类别特定的追踪器使用模式。
  • 对超过 9 万个域名按主题(如健康、成瘾)进行分类,比较不同内容类型之间的追踪器使用率。
  • 将追踪器分布与国家层面的政治指标(如新闻自由和美国政治立场)进行相关性分析。

实验结果

研究问题

  • RQ1哪些第三方追踪服务在网页中最为普遍,它们分别追踪了多少个网站?
  • RQ2网页追踪器的分布如何在不同国家之间变化,政治或社会文化因素在其中扮演什么角色?
  • RQ3涵盖隐私敏感主题(如健康或成瘾)的网站是否比其他网站使用更少的追踪器?
  • RQ4追踪器主导地位在多大程度上由地缘政治因素驱动,而非经济或技术因素?

主要发现

  • 谷歌、脸书和推特是三大最占主导地位的追踪服务,且所有三个排名靠前的 PageRank 追踪器域名均由谷歌拥有。
  • 追踪网络在所追踪网站数量上表现出幂律分布,表明少数大型追踪器占据主导地位。
  • 追踪网络具有异配性,即高阶追踪器(主要平台)倾向于连接到低阶追踪器,表明存在集中化控制。
  • 尽管谷歌官方已从中国退出,其追踪服务在中文网站上仍保持活跃,表明其持续的业务存在。
  • 中国、俄罗斯和伊朗等国家展现出独特的追踪器使用模式,主要受政治因素(如新闻自由和美国政治立场)影响,而非经济指标。
  • 尽管隐私敏感主题网站(如健康与成瘾)的追踪器使用率为 60%,显著低于普通网站的 90%,但整体水平仍然偏高。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。