[Paper Review] On the Ubiquity of Web Tracking: Insights from a Billion-Page Web Crawl
This study analyzes third-party web tracking using a billion-page crawl from CommonCrawl 2012, extracting 140 million tracker embeddings across 41 million domains. It reveals that Google, Facebook, and Twitter dominate tracking globally, with political factors like press freedom explaining deviations—especially in China and Russia—while privacy-critical sites still show high tracker prevalence (60%) despite lower rates than general sites.
We perform a large-scale analysis of third-party trackers on the World Wide Web from more than 3.5 billion web pages of the CommonCrawl 2012 corpus. We extract a dataset containing more than 140 million third-party embeddings in over 41 million domains. To the best of our knowledge, this constitutes the largest web tracking dataset collected so far, and exceeds related studies by more than an order of magnitude in the number of domains and web pages analyzed. We perform a large-scale study of online tracking, on three levels: (1) On a global level, we give a precise figure for the extent of tracking, give insights into the structure of the `online tracking sphere' and analyse which trackers are used by how many websites. (2) On a country-specific level, we analyse which trackers are used by websites in different countries, and identify the countries in which websites choose significantly different trackers than in the rest of the world. (3) We answer the question whether the content of websites influences the choice of trackers they use, leveraging more than 90 thousand categorized domains. In particular, we analyse whether highly privacy-critical websites make different choices of trackers than other websites. Based on the performed analyses, we confirm that trackers are widespread (as expected), and that a small number of trackers dominates the web (Google, Facebook and Twitter). In particular, the three tracking domains with the highest PageRank are all owned by Google. The only exception to this pattern are a few countries such as China and Russia. Our results suggest that this dominance is strongly associated with country-specific political factors such as freedom of the press. We also confirm that websites with highly privacy-critical content are less likely to contain trackers (60% vs 90% for other websites), even though the majority of them still do contain trackers.
Motivation & Objective
- To understand the global scale and structural properties of third-party web tracking using a massive dataset.
- To investigate how tracker usage varies across countries and whether political, cultural, or economic factors influence these patterns.
- To examine whether websites covering privacy-sensitive topics like health and addiction use fewer trackers than other sites.
- To identify dominant tracking services and assess their reach across the web, especially in relation to geopolitical factors.
Proposed method
- Extracted third-party tracker embeddings from 3.5 billion web pages in the CommonCrawl 2012 corpus, focusing on pay-level domains (PLDs).
- Constructed a bipartite tracking graph linking websites (PLDs) to third-party tracking services, enabling large-scale network analysis.
- Applied network analysis techniques including PageRank, degree distribution, and dissortativity to study structural properties of the tracking sphere.
- Used clustering of co-occurrence patterns to identify country-specific and category-specific tracker usage patterns.
- Categorized over 90,000 domains by topic (e.g., health, addiction) to compare tracker prevalence across content types.
- Correlated tracker distribution with country-level political indicators such as press freedom and US political alignment.
Experimental results
Research questions
- RQ1Which third-party tracking services are most prevalent across the web, and how many websites do they track?
- RQ2How does the distribution of web trackers vary across different countries, and what role do political or socio-cultural factors play?
- RQ3Do websites covering privacy-critical topics such as health or addiction use fewer trackers than other websites?
- RQ4To what extent is tracker dominance driven by geopolitical factors rather than economic or technical factors?
Key findings
- Google, Facebook, and Twitter are the three most dominant tracking services, with all three top PageRank tracker domains owned by Google.
- The tracking network exhibits a power-law distribution in the number of websites tracked, indicating a few large trackers dominate.
- The tracking network is dissortative, meaning high-degree trackers (major platforms) tend to link to low-degree ones, suggesting centralized control.
- Despite Google’s official retreat from China, its tracking services remain active on Chinese websites, indicating continued operational presence.
- Countries like China, Russia, and Iran show distinct tracker usage patterns, primarily due to political factors such as press freedom and US political alignment, not economic indicators.
- Websites covering privacy-critical topics like health and addiction still use trackers in 60% of cases, significantly lower than the 90% rate on general websites, but still high overall.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.