Skip to main content
QUICK REVIEW

[論文レビュー] On the Ubiquity of Web Tracking: Insights from a Billion-Page Web Crawl

Sebastian Schelter, Jérôme Kunegis|arXiv (Cornell University)|Jul 25, 2016
Privacy, Security, and Data Protection参考文献 25被引用数 4
ひとこと要約

本研究では、CommonCrawl 2012の10億ページにわたるクロールデータを用いて、第三者ウェブトラッキングを分析し、4100万ドメインにまたがる1億4000万のトラッカー埋め込みを抽出した。その結果、グーグル、フェイスブック、ツイッターが世界規模でトラッキングを支配していることが明らかになった。政治的要因、特にジャーナリズムの自由度が中国やロシアなどでの乖離を説明しており、プライバシーに配慮するサイトでもトラッカーの使用率は依然として60%に達する。

ABSTRACT

We perform a large-scale analysis of third-party trackers on the World Wide Web from more than 3.5 billion web pages of the CommonCrawl 2012 corpus. We extract a dataset containing more than 140 million third-party embeddings in over 41 million domains. To the best of our knowledge, this constitutes the largest web tracking dataset collected so far, and exceeds related studies by more than an order of magnitude in the number of domains and web pages analyzed. We perform a large-scale study of online tracking, on three levels: (1) On a global level, we give a precise figure for the extent of tracking, give insights into the structure of the `online tracking sphere' and analyse which trackers are used by how many websites. (2) On a country-specific level, we analyse which trackers are used by websites in different countries, and identify the countries in which websites choose significantly different trackers than in the rest of the world. (3) We answer the question whether the content of websites influences the choice of trackers they use, leveraging more than 90 thousand categorized domains. In particular, we analyse whether highly privacy-critical websites make different choices of trackers than other websites. Based on the performed analyses, we confirm that trackers are widespread (as expected), and that a small number of trackers dominates the web (Google, Facebook and Twitter). In particular, the three tracking domains with the highest PageRank are all owned by Google. The only exception to this pattern are a few countries such as China and Russia. Our results suggest that this dominance is strongly associated with country-specific political factors such as freedom of the press. We also confirm that websites with highly privacy-critical content are less likely to contain trackers (60% vs 90% for other websites), even though the majority of them still do contain trackers.

研究の動機と目的

  • 大規模なデータセットを用いて、第三者ウェブトラッキングの世界的な規模と構造的特性を理解すること。
  • トラッキングの使用状況が国ごとにどのように異なるか、また政治的・文化的・経済的要因がそのパターンに与える影響を調査すること。
  • 健康や依存症といったプライバシーに敏感なトピックを扱うウェブサイトが、他のサイトと比較してトラッカーを少なく使用しているかどうかを検討すること。
  • 支配的トラッキングサービスを特定し、特に地政学的要因に関連してそのウェブ全体への浸透度を評価すること。

提案手法

  • CommonCrawl 2012コーパスに含まれる35億ページから、ペイロードレベルドメイン(PLD)を対象に、第三者トラッキング埋め込みを抽出した。
  • ウェブサイト(PLD)と第三者トラッキングサービスを結ぶ二部グラフを構築し、大規模なネットワーク解析を可能にした。
  • ページランク、次数分布、非同一性(dissortativity)といったネットワーク解析手法を用いて、トラッキング領域の構造的特性を分析した。
  • 共起パターンのクラスタリングを用いて、国別およびカテゴリ別のトラッキング使用パターンを同定した。
  • 健康や依存症など9万件を超えるドメインをトピック別(例:健康、依存症)に分類し、コンテンツタイプごとのトラッキング頻度を比較した。
  • ジャーナリズムの自由度や米国の政治的傾向といった国レベルの政治指標と、トラッキングの分布を相関させた。

実験結果

リサーチクエスチョン

  • RQ1ウェブ全体で最も一般的な第三者トラッキングサービスは何か。また、それらはどれくらいのウェブサイトを追跡しているか。
  • RQ2ウェブトラッキングの分布は国ごとにどのように異なるのか。政治的または社会文化的要因はその要因として果たす役割は何か。
  • RQ3健康や依存症といったプライバシーに敏感なトピックを扱うウェブサイトは、他のサイトと比較してトラッキングを少なく使用しているのか。
  • RQ4トラッキング支配の背後には、経済的要因や技術的要因よりも、地政学的要因がどれほど影響を与えているのか。

主な発見

  • グーグル、フェイスブック、ツイッターが、最も支配的なトラッキングサービスであり、上位3つのページランクを持つトラッキングドメインはすべてグーグル所有である。
  • トラッキングネットワークは、追跡対象となるウェブサイト数に指数分布(パワーロウ)を示しており、少数の大きなトラッカーが支配的であることが示された。
  • トラッキングネットワークは非同一性(dissortative)である。つまり、次数の高いトラッカー(主要プラットフォーム)は、次数の低いトラッカーに接続しており、中央集権的制御が特徴的である。
  • グーグルが中国での事業を公式に撤退したにもかかわらず、中国のウェブサイトでは依然としてそのトラッキングサービスが活用されており、継続的な存在が裏付けられた。
  • 中国、ロシア、イランのような国々では、ジャーナリズムの自由度や米国の政治的傾向といった政治的要因が、主な要因となって、特異なトラッキング使用パターンを示している。経済的指標はその要因とはなっていない。
  • 健康や依存症といったプライバシーに配慮するトピックを扱うウェブサイトでも、トラッキングの使用率は依然として60%に達しており、一般サイトの90%とは比較して低いが、依然として高い水準にある。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。